# Website Contact Scraper: Emails, Phones & Socials, Fixed Price (`accountable_eel/website-contact-extractor`) Actor

Contact info scraper alternative: a fixed price per website and a hard page cap, so no runaway crawls. For a list of websites, get emails (role inboxes flagged), E.164 phones, LinkedIn, X, Facebook, Instagram, YouTube, TikTok, contact form and address. Export to Google Sheets. Empty sites are free.

- **URL**: https://apify.com/accountable\_eel/website-contact-extractor.md
- **Developed by:** [Adrian Voss](https://apify.com/accountable_eel) (community)
- **Categories:** Lead generation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 website with results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scraper: Emails, Phones & Socials, Fixed Price per Site

**One website in, one row out, one flat price, and no runaway crawls.** Give it a list of websites
and get back, per site: every public email (with shared inboxes like info@ and sales@ flagged),
phone numbers in international E.164 format, the LinkedIn company page, X, Facebook, Instagram,
YouTube and TikTok profiles, the contact form page and the postal address. Ready to export to
Google Sheets, a CRM or Clay.

It is a contact info scraper alternative built around one complaint about the popular ones: you
pay per page crawled, and a crawl that wanders into a blog or a shop can cost fifty times what
you expected. Here the price is per **site**, whatever the site's size, and every site has a
**hard page cap** (8 by default, 20 at most). A site where nothing is found is free.

### Who it's for

- **Sales and lead-gen teams** enriching a list of websites (from a trade-show list, Google Maps
  export or CRM) with a general inbox, phone number and LinkedIn page before outreach.
- **Agencies and freelancers** building prospect lists for clients who need a predictable bill:
  1,000 sites costs the same whether they are one-pagers or 10,000-page corporate sites.
- **Clay, n8n and Make users** who want one clean row per domain, flat columns (`primaryEmail`,
  `primaryPhone`, `linkedinUrl`) that map straight into a table.
- **Researchers and data teams** who need the legal-notice address of European companies (the
  German/Austrian/Swiss *Impressum*, French *mentions légales*) without reading each site.

### Why this one

- **Fixed price per website.** One charge per site with results, whether it took 1 page or 20.
  A site that returns nothing, can't be reached or blocks the request costs nothing.
- **No runaway crawls.** It reads the homepage plus **only** the links that look like contact
  pages: contact, kontakt, about, team, impressum, imprint, legal, mentions légales, support,
  privacy. Blog posts, product pages and shops are never opened. `maxPagesPerSite` is a hard cap.
- **Emails you can use.** Deduplicated, lower-cased, each with the page it was found on
  (`sourcePage`), `isRoleAddress` for shared inboxes and `onSiteDomain` for addresses on the
  company's own domain. Cloudflare-protected emails (`[email protected]`) and simple
  `name [at] company [dot] com` forms are decoded. Placeholders like `name@company.com` are dropped.
- **Phones in E.164.** `+4998315050`, not "09831 505-0". Numbers without a country prefix are read
  using the site's country (from its address, its .de/.fr/.co.uk domain, or its page language).
  Fax numbers are left out.
- **Respects robots.txt.** robots.txt is read before any page, and disallowed pages are skipped
  (and counted in `pagesSkippedByRobots`).
- **Plain HTTP, no browser.** Fast and cheap. The trade-off: contact details that only appear
  after JavaScript runs are not seen (see FAQ).

### Price

- **Website with results**: $20 per 1,000 websites

Plus a $0.00005 start fee per run. Each event above is billed independently, only when it actually returns data — misses (`found:false`) are never charged.

The price per site does not depend on `maxPagesPerSite`: raising the cap to 20 reads more pages
for the same charge.

### How to use

1. **In the Apify Console.** Open the actor page and click **Start** — the `websites` field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
2. **Via the API.** Call it directly with a POST request — no Console needed once you have an API token:
   ```bash
   curl "https://api.apify.com/v2/acts/accountable_eel~website-contact-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
     -X POST \
     -H "Content-Type: application/json" \
     -d '{"websites":["hetzner.com","fastly.com"]}'
   ```
3. **On a schedule.** Save this actor as an Apify **Task** with the input you want, then add a **Schedule** (hourly, daily, weekly) so it runs on its own — no server of your own required.

**For a list of websites in Google Sheets:** paste the domain column into **Websites** (one per
line; `acme.com`, `www.acme.com` and full URLs all work), run, then export the dataset as CSV or
connect the Google Sheets integration. Use the flat columns `primaryEmail`, `emailsText`,
`primaryPhone`, `phonesText`, `linkedinUrl` and `addressText`: one cell each, no nested JSON.

### Input

```json
{
  "websites": [
    "hetzner.com",
    "fastly.com"
  ]
}
```

One website per line: acme.com, www.acme.com or https://acme.com/en/ all work. Each site is read as the homepage plus only its contact-looking pages (contact, impressum, about, team, legal, support, privacy), capped by 'Max pages per site' below. One flat price per site with results; a site where nothing is found is free. Accepted formats: hetzner.com, www.fastly.com, https://earlydesigneducation.gsd.harvard.edu/.

### Output

One row per website. A site with results looks like this (trimmed):

```json
{
  "query": "https://hetzner.com/",
  "found": true,
  "domain": "hetzner.com",
  "website": "https://www.hetzner.com/",
  "primaryEmail": "info@hetzner.com",
  "emailsText": "info@hetzner.com, tco@hetzner.com, dsa-authorities@hetzner.com",
  "emails": [
    { "email": "info@hetzner.com", "isRoleAddress": true, "onSiteDomain": true, "sourcePage": "https://www.hetzner.com/legal/legal-notice/" }
  ],
  "primaryPhone": "+4998315050",
  "phones": [
    { "number": "+49 9831 505-0", "e164": "+4998315050", "country": "DE", "sourcePage": "https://www.hetzner.com/legal/legal-notice/" }
  ],
  "linkedinUrl": "https://www.linkedin.com/company/hetzner-online",
  "xUrl": "https://x.com/Hetzner_Online",
  "facebookUrl": "https://www.facebook.com/hetzner.de",
  "instagramUrl": "https://www.instagram.com/hetzner.online",
  "youtubeUrl": "https://www.youtube.com/user/HetznerOnline",
  "contactFormUrl": null,
  "addressText": "Industriestr. 25, 91710 Gunzenhausen, DE",
  "address": { "streetAddress": "Industriestr. 25", "postalCode": "91710", "addressLocality": "Gunzenhausen", "addressCountry": "DE", "source": "impressum" },
  "pagesFetched": 5,
  "pageCapReached": true
}
```

| query | found | status | domain | website | primaryEmail | emailsText | emails | emailCount | primaryPhone | phonesText | phones | phoneCount | phoneCountry | linkedinUrl | xUrl | facebookUrl | instagramUrl | youtubeUrl | tiktokUrl | otherSocialUrls | contactFormUrl | addressText | address | pagesFetched | pagesRead | pagesSkippedByRobots | pageCapReached | scrapedAt |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| hetzner.com | true | OK | <domain> | <homepage read> | <main email> | \<all emails (comma-separated)> | \<emails (with source page)> | <emails found> | <main phone> | \<all phones (comma-separated)> | \<phones (with source page)> | <phones found> | <country used for local numbers> | <linkedin company page> | \<x (twitter)> | <facebook> | <instagram> | <youtube> | <tiktok> | <other social profiles> | <contact form page> | <postal address> | \<postal address (parts)> | <pages fetched> | <pages read> | \<pages skipped (robots.txt)> | <page cap reached> | 1970-01-01T00:00:00.000Z |

A site with nothing found, an unreachable site, or one whose robots.txt disallows it comes back as
a row with `"found": false`, a `status` (`NOT_FOUND`, `BLOCKED`, `REQUEST_FAILED`, `BAD_FORMAT`)
and a plain-English `message`. Those rows are never charged. Turn on "Only return found results"
to hide them.

### Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

```bash
curl "https://api.apify.com/v2/acts/accountable_eel~website-contact-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -X POST \
  -H "Content-Type: application/json" \
  -d '{"websites":["hetzner.com","fastly.com"]}'
```

**n8n.** Add an HTTP Request node: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~website-contact-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body Content Type `JSON`, JSON Body `{"websites":["hetzner.com","fastly.com"]}` (swap in an expression from an earlier node for a real value).

**Clay.** Add an "HTTP API" column: Method `POST`, URL `https://api.apify.com/v2/acts/accountable_eel~website-contact-extractor/run-sync-get-dataset-items?token=<YOUR_TOKEN>`, Body `{"websites":["{{website}}"]}`, mapping the row's website into the `websites` array.

**MCP.** In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "Website Contact Scraper: Emails & Phones, Fixed Price" — the agent will find and run this actor.

### vs. alternatives

| | What it costs | What you get | Trade-off |
|---|---|---|---|
| **This actor** | $20 per 1,000 websites with results (less on paid plans), nothing for sites with no results | Emails with source page and role flag, E.164 phones, 6 social networks, contact form, postal address, up to 20 contact pages per site | Plain HTTP: details that only render with JavaScript are missed |
| Popular contact detail scrapers on the Apify Store | Per page crawled: about $2 per 1,000 pages on the free plan, so an 8-page site is ~$0.016 and a deep crawl much more | Crawls to a chosen depth, optional browser rendering, paid lead enrichment add-ons | The bill depends on how big each site is; a missing depth limit is the classic surprise |
| Email finder tools (Hunter-style) | Monthly subscription or credits | Guessed and verified personal work emails | Guesses rather than what the site publishes; no phones, socials or address |
| Doing it by hand | Your time | Exactly what you look for | Minutes per site |

Store figures as of September 2026.

### Data & privacy

This actor reads only public pages a visitor can open without logging in, and it honours each
site's robots.txt. It returns the contact details the website itself publishes. It does not guess
email addresses, look anyone up in third-party databases or enrich people with other sources. Some
sites publish named staff emails and phone numbers; those are personal data under the GDPR and
similar laws, so make sure you have a lawful basis (for example legitimate interest for B2B
outreach, with an opt-out) before you use them, and follow the anti-spam rules where you send.

### FAQ

**How is "fixed price per site" different from per-result pricing?**
You pay one site event per website that returned at least one email, phone, social profile,
contact form or address, however many pages were read and however many emails it had. Sites
with nothing, unreachable sites and blocked sites are free.

**What does `maxPagesPerSite` do, and does it change the price?**
It is the most pages that will ever be fetched for one site, homepage included (1 to 20, default
8\). It does not change the price. Most sites have 2 to 6 contact-looking pages, so the default
rarely bites. If `pageCapReached` is `true`, there were more such pages than your cap.

**Which pages does it read?**
The homepage (or the URL you gave), then links on it whose address or text looks like a contact
page, in this order: contact, legal notice / impressum, about / team, support, privacy. Only the
same host (www and the bare domain count as one); turn on "Follow links to subdomains" to include
e.g. support.acme.com. If the homepage has no such links (a site whose menu is built by
JavaScript), it tries `/contact` and `/about` (or `/impressum` on .de/.at/.ch domains).

**Why is an email or phone I can see in my browser missing?**
Either it is on a page the crawler didn't open (raise the cap), or the site draws it with
JavaScript, which a plain HTTP reader never sees. Emails shown only as an image are also missed.

**Does it guess emails like firstname.lastname@?**
No. It returns only what the site publishes. To guess and check a person's work email from their
name and the company domain, use [Email Finder by Name and Domain](https://apify.com/accountable_eel/email-pattern-finder).

**What is `isRoleAddress`?**
`true` for shared inboxes (info@, sales@, support@, contact@, hello@, kontakt@, press@ and
similar), `false` for addresses that look like a person's. Good for sorting a general inbox from
a named contact.

**Why is `e164` empty for some phone numbers?**
The number was written without a country code and the site's country couldn't be told from its
address, domain or page language. `number` keeps it exactly as written.

**Does it respect robots.txt?**
Yes. robots.txt is read first; a site that disallows its homepage is not crawled at all (a free
`BLOCKED` row), and disallowed contact pages are skipped.

**Can an AI agent call this?**
Yes, through the Apify MCP server or the API call shown above. Ask for "Website Contact Scraper:
Emails, Phones & Socials".

### Related actors

- [Email Finder by Name and Domain](https://apify.com/accountable_eel/email-pattern-finder): the
  company's email format and a ranked guess for a named person.
- [Company Domain Enrichment](https://apify.com/accountable_eel/company-domain-enrichment): DNS,
  email provider, tech stack, registration and hiring signals for the same list of domains.
- [Email Deliverability Check](https://apify.com/accountable_eel/email-deliverability-check): check
  the emails this actor found for a mail route, disposable domains and catch-all.
- [JSON-LD Structured Data Lookup](https://apify.com/accountable_eel/jsonld-structured-data-lookup):
  all schema.org data a page publishes.

# Actor input Schema

## `websites` (type: `array`):

One website per line: acme.com, www.acme.com or https://acme.com/en/ all work. Each site is read as the homepage plus only its contact-looking pages (contact, impressum, about, team, legal, support, privacy), capped by 'Max pages per site' below. One flat price per site with results; a site where nothing is found is free. Accepted formats: hetzner.com, www.fastly.com, https://earlydesigneducation.gsd.harvard.edu/. You're only charged for the ones we actually find — a miss costs nothing.

## `testRun` (type: `boolean`):

Turn this on to test your input on a small sample before running the full list. Turn it off to process everything.

## `onlyFound` (type: `boolean`):

Only keep rows where something was actually found. Misses are always free, whether or not you show them here.

## `includeKeywords` (type: `array`):

Optional. Only keep results that mention at least one of these words (e.g. a job title, a city, a product name). Leave empty to keep everything.

## `excludeKeywords` (type: `array`):

Optional. Drop any result that mentions one of these words. Leave empty to skip nothing.

## `maxResults` (type: `integer`):

Optional. Stop the run once this many results have been found — useful for a quick, cheap sample. Leave blank for no limit.

## `columns` (type: `array`):

Choose which pieces of information to include in each result row. All are included by default.

## `maxPagesPerSite` (type: `integer`):

1 to 20, homepage included. Only the homepage and links that look like contact pages are ever fetched, so most sites stop well before the cap. The price per site is the same whatever this is set to.

## `includeSubdomains` (type: `boolean`):

Off by default: only the site's own host is read (www. and the bare domain count as the same host). Turn on to also follow contact links to e.g. support.acme.com. Still inside the page cap.

## `maxConcurrency` (type: `integer`):

Parallel requests. Keep conservative — this target has no browser fallback, so getting blocked costs more than slow-and-steady.

## `proxyConfiguration` (type: `object`):

Apify Proxy config. Residential recommended for anti-bot-sensitive targets.

## Actor input object example

```json
{
  "websites": [
    "hetzner.com",
    "fastly.com"
  ],
  "testRun": false,
  "onlyFound": false,
  "includeKeywords": [],
  "excludeKeywords": [],
  "columns": [
    "domain",
    "website",
    "primaryEmail",
    "emailsText",
    "emails",
    "emailCount",
    "primaryPhone",
    "phonesText",
    "phones",
    "phoneCount",
    "phoneCountry",
    "linkedinUrl",
    "xUrl",
    "facebookUrl",
    "instagramUrl",
    "youtubeUrl",
    "tiktokUrl",
    "otherSocialUrls",
    "contactFormUrl",
    "addressText",
    "address",
    "pagesFetched",
    "pagesRead",
    "pagesSkippedByRobots",
    "pageCapReached"
  ],
  "maxPagesPerSite": 8,
  "includeSubdomains": false,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "hetzner.com",
        "fastly.com"
    ],
    "includeKeywords": [],
    "excludeKeywords": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("accountable_eel/website-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "hetzner.com",
        "fastly.com",
    ],
    "includeKeywords": [],
    "excludeKeywords": [],
}

# Run the Actor and wait for it to finish
run = client.actor("accountable_eel/website-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "hetzner.com",
    "fastly.com"
  ],
  "includeKeywords": [],
  "excludeKeywords": []
}' |
apify call accountable_eel/website-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,accountable_eel/website-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qDQemIDGVUNux3ocJ/builds/yCmscBUYa0QH0vBGc/openapi.json
