# Website Contact Finder – Emails, Phones & Social Links (`gazidev/website-contact-finder`) Actor

Find emails, phone numbers and social media profiles (LinkedIn, X/Twitter, Facebook, Instagram, YouTube, TikTok) on any list of company websites. Crawls contact, about, impressum and legal pages, one clean row per domain. Pay only for domains where contacts are found.

- **URL**: https://apify.com/gazidev/website-contact-finder.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 website with contacts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Finder – Emails, Phones & Social Links

Give it a list of company websites. For each one you get back **email addresses, phone numbers and social media profiles** (LinkedIn, X/Twitter, Facebook, Instagram, YouTube, TikTok, GitHub, Pinterest), plus the **contact page URL, company name, address and logo**. Each domain comes back as **one clean, deduplicated row**, ready for your CRM or spreadsheet.

- **Pay per domain, only when contacts are found.** You pay $4 per 1,000 domains that return at least one email, phone or social profile. Domains with no contacts, dead domains, blocked sites and errors cost nothing.
- **Smart page picking.** The Actor crawls the homepage, then the pages most likely to hold contact details: *Contact, About, Impressum/Imprint, Team, Legal notice, Locations* and footer links. It understands 15+ languages (kontakt, contacto, contatti, iletişim, mentions légales, over ons…). It does not crawl blog posts or product pages.
- **Clean emails.** It finds `mailto:` links, plain text, obfuscated forms like `name [at] domain [dot] com` and `name (at) domain.com`, Cloudflare-protected addresses and schema.org data. Junk is filtered out: image names (`logo@2x.png`), Sentry/Wix tracking addresses, placeholders (`name@example.com`) and third-party widget addresses. A **DNS MX check** tells you whether each email domain can actually receive mail.
- **Normalized phone numbers.** Numbers come from `tel:` links, schema.org and page text, parsed with Google's libphonenumber. Each one is returned in international and E.164 format with its country and line type (mobile, landline, toll-free). **Fax numbers are separated out.** The phone country is detected per site from the country domain and the page language.
- **Real social profiles only.** Share buttons, intent links, individual posts and videos are ignored. You get the company's actual profile URLs, normalized and deduplicated.
- **Polite and fast.** HTTP only (no browser) and **robots.txt respected**, including Crawl-delay. Each site gets at most 2 parallel requests, and about 80–100 sites are processed per minute at the default settings.

### What can I use it for?

- **Lead generation and sales prospecting:** turn a list of company domains (from Google Maps, a trade-fair exhibitor list, a CRM export or a directory) into emails, phone numbers and LinkedIn pages
- **CRM enrichment:** add missing contact data and social profiles to existing accounts
- **Agencies and freelancers:** build local business lists (restaurants, dentists, shops, hotels) with their contact details
- **Partnership and PR outreach:** find the press, sales or partnership emails of brands
- **Data quality:** check that companies still list the same contact details, and flag email domains that no longer receive mail
- **AI agents:** give an agent a "find contact info for this company" tool (see below)

### Input

| Field | Description |
|---|---|
| `urls` | Domains or URLs (`acme.com`, `https://www.acme.de/en/`). Each domain is processed once. |
| `bulkText` | Paste any list: one per line, comma separated, or a whole CSV export |
| `sourceFileUrl` | Public URL of a `.txt`/`.csv` file (e.g. a published Google Sheet) |
| `maxPagesPerSite` | Pages per website, homepage included (default 10, max 50). **The price per domain stays the same.** |
| `maxDepth` | 0 = homepage only, 1 = + contact-like pages linked from it, 2 = + one more level (default) |
| `extractEmails`, `extractPhones`, `extractSocials` | Choose what to find |
| `verifyEmailDomains` | DNS MX check for every email domain (default on) |
| `phoneCountry` | Force a country for local-format numbers (`DE`, `GB`…). Empty = auto per site |
| `includePageDetails` | Add a `pages` array showing what was found on each crawled URL |
| `maxSites`, `respectRobotsTxt`, `maxConcurrency`, `requestTimeoutSecs`, `proxyConfiguration` | Advanced |

```json
{
  "urls": ["hawksheadrelish.com", "puccinibomboni.com", "fairmountbagel.com"],
  "maxPagesPerSite": 10,
  "maxDepth": 2,
  "extractEmails": true,
  "extractPhones": true,
  "extractSocials": true,
  "verifyEmailDomains": true
}
```

### Output

You get one row per domain. The **Contacts** and **Social profiles** tables in the Output tab show the main columns; the full JSON looks like this (shortened):

```json
{
  "domain": "puccinibomboni.com",
  "url": "https://puccinibomboni.com",
  "companyName": "Puccini Bomboni",
  "emails": ["customerservice@puccinibomboni.com", "info@puccinibomboni.com", "webshop@puccinibomboni.com"],
  "phones": ["+31 20 626 5474", "+31 20 427 8341", "+31 20 737 0530"],
  "faxNumbers": [],
  "linkedin": "https://www.linkedin.com/company/puccini-bomboni",
  "facebook": "https://www.facebook.com/PucciniBomboniAmsterdam",
  "instagram": "https://www.instagram.com/puccinibomboni_amsterdam",
  "twitter": null, "youtube": null, "tiktok": null, "github": null, "pinterest": null,
  "socialLinks": {"linkedin": ["..."], "facebook": ["..."], "instagram": ["..."]},
  "contactPageUrl": "https://puccinibomboni.com/contact",
  "hasContactForm": true,
  "emailDetails": [
    {"email": "customerservice@puccinibomboni.com", "domainMatchesSite": true, "mxValid": true,
     "foundOnPages": ["https://puccinibomboni.com/faq"], "occurrences": 2}
  ],
  "phoneDetails": [
    {"phone": "+31 20 626 5474", "e164": "+31206265474", "country": "NL", "type": "fixed_line",
     "isFax": false, "source": "tel-link", "foundOnPages": ["https://puccinibomboni.com/contact"], "occurrences": 20}
  ],
  "address": null,
  "language": "nl-NL",
  "contactsFound": 9,
  "pagesCrawled": 9,
  "robotsTxt": "ok",
  "jsRenderedSite": false,
  "blocked": false,
  "error": null
}
```

Emails on the company's own domain come first. `domainMatchesSite: false` marks outside addresses such as a Gmail inbox, a parent company or the web agency credited in the footer. Sites that fail have `error` set (`DNS resolution failed`, `Blocked by bot protection`, `Disallowed by robots.txt`…), and **those rows are free**.

See [SAMPLE_OUTPUT.json](SAMPLE_OUTPUT.json) for complete real results.

### Pricing

Pay per event:

| Event | Price |
|---|---|
| Domain with contacts (at least one email, phone or social profile found) | **$0.004** ($4 per 1,000 domains) |
| Actor start | $0.0005 |
| Domains with no contacts, failed or blocked domains | free |

How that compares with other contact scrapers in the Apify Store (listed prices, September 2026):

| Actor | Listed price | Approx. per 1,000 domains\* |
|---|---|---|
| vdrmota/contact-info-scraper (market leader) | $2 per 1,000 **pages** | $10–20 (at 5–10 pages per domain) |
| emastra | $4 per 1,000 results | about $4+ |
| jurassic_jove | $6 per 1,000 results | about $6+ |
| delicious_zebu | $10 per 1,000 results | about $10+ |
| caprolok | $20 per 1,000 results | about $20+ |
| **This Actor** | **$4 per 1,000 domains with contacts** | **$4, and $0 for domains without contacts** |

\*Our estimate. Prices differ in unit (pages, results or domains), so check each Actor's pricing tab.

Set a **maximum cost per run** in the run options. The Actor stops cleanly once the next domain would go over it.

### FAQ

**Is it legal to collect contact data from websites?**
This Actor only reads **publicly published business contact information** (the kind of details companies put on their contact, imprint and legal pages so that people can reach them). It respects robots.txt. You are responsible for how you use the data. Personal data such as a named employee's email is still personal data under the **GDPR/UK GDPR**, and email marketing is regulated (for example **CAN-SPAM**, **ePrivacy/PECR** and **CASL**). Make sure you have a lawful basis (such as legitimate interest for B2B outreach), honor opt-outs, and do not use the data for spam. If you are unsure, ask a lawyer. Apify's article [Is web scraping legal?](https://blog.apify.com/is-web-scraping-legal/) is a good starting point.

**Why did a site return no contacts?**
Common reasons:

- The site shows contact details only inside a JavaScript app (`jsRenderedSite: true`) or only through a contact form (`hasContactForm: true`).
- The site blocks automated HTTP clients (`blocked: true`).
- The domain is parked or for sale.

You are not charged for these domains.

**Does it use a headless browser?**
No. It uses plain HTTP, which is what keeps it fast and cheap. Sites protected by Cloudflare/Incapsula "checking your browser" pages are reported as `blocked` rather than retried with an expensive browser.

**How are pages chosen?**
Links on the homepage are scored by URL and link text: contact/imprint pages first, then about/team/legal/locations, then privacy/terms. Footer links get a bonus. If no contact link exists, the Actor tries common paths such as `/contact` and `/impressum`. Other subdomains and sister domains of the same brand are included, and all other websites are ignored.

**Can I get the page where each email was found?**
Yes. `emailDetails[].foundOnPages` and `phoneDetails[].foundOnPages` list the URLs, and `includePageDetails: true` adds a per-page breakdown.

**Can I feed it a Google Sheet?**
Yes. Publish the sheet as CSV and paste the link into `sourceFileUrl`. To send the results back to a sheet, use a Google Sheets integration.

### Use with AI agents / Apify MCP

The Actor works well as a tool for AI agents. It takes a small input (a domain list), returns one compact row per domain and costs very little. Through the [Apify MCP server](https://mcp.apify.com) (`https://mcp.apify.com?actors=gazidev/website-contact-finder`), Claude, ChatGPT, Cursor and other MCP clients can call it directly, for example: *"find the contact email and LinkedIn page of these 20 companies"*. You can also call it over the Apify API: `POST https://api.apify.com/v2/acts/gazidev~website-contact-finder/run-sync-get-dataset-items` with the input JSON.

### Categories

Lead generation · Marketing · SEO tools · Automation · Developer tools

# Actor input Schema

## `urls` (type: `array`):

Company websites to scan, e.g. `acme.com` or `https://www.acme.de/en/`. `https://` is added automatically. Each domain is processed once and returns one row.

## `bulkText` (type: `string`):

Paste a big list: one domain per line, comma/semicolon separated, or a whole CSV export (the first cell that looks like a domain/URL in each row is used).

## `sourceFileUrl` (type: `string`):

Public URL of a .txt or .csv file with domains/URLs (e.g. a Google Sheets 'publish to web' CSV link).

## `maxPagesPerSite` (type: `integer`):

Homepage + the most promising pages (contact, about, impressum/imprint, team, legal, footer links). 10 is enough for most small business sites. The price per domain is the same whatever you choose.

## `maxDepth` (type: `integer`):

0 = homepage only, 1 = homepage + contact-like pages linked from it, 2 = also contact-like pages linked from those (e.g. Contact → Locations).

## `extractEmails` (type: `boolean`):

Find email addresses (mailto links, text, obfuscated `name [at] domain [dot] com`, Cloudflare-protected emails, schema.org). Junk (image names, Sentry/Wix tracking addresses, placeholders) is filtered out.

## `extractPhones` (type: `boolean`):

Find phone numbers (tel: links, schema.org, text) and normalize them with Google's libphonenumber (international + E.164 format, country, mobile/landline, fax flagged).

## `extractSocials` (type: `boolean`):

LinkedIn, X/Twitter, Facebook, Instagram, YouTube, TikTok, GitHub and Pinterest profile links (share buttons and post links are ignored).

## `verifyEmailDomains` (type: `boolean`):

Check that each email's domain can receive mail (MX record). Emails on dead domains are moved out of `emails` (still listed in `emailDetails` with `mxValid: false`).

## `phoneCountry` (type: `string`):

Two-letter country code used to read local-format numbers like `030 1234567`, e.g. `DE`, `GB`, `US`. Leave empty to detect it per site from the country domain (.de, .co.uk...) and page language.

## `includePageDetails` (type: `boolean`):

Add a `pages` array to every row: each crawled URL with the emails, phones and social links found on it.

## `maxSites` (type: `integer`):

Stop after this many websites (0 = no limit). Useful to test a big list cheaply.

## `respectRobotsTxt` (type: `boolean`):

Skip pages disallowed by the site's robots.txt and honor Crawl-delay (max 5 s). Recommended.

## `maxConcurrency` (type: `integer`):

How many websites are crawled at the same time. Each website gets at most 2 parallel requests, so crawling stays polite.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout. Failed requests are retried.

## `proxyConfiguration` (type: `object`):

Optional. Not needed for most sites.

## Actor input object example

```json
{
  "urls": [
    "hawksheadrelish.com",
    "puccinibomboni.com",
    "fairmountbagel.com"
  ],
  "maxPagesPerSite": 10,
  "maxDepth": 2,
  "extractEmails": true,
  "extractPhones": true,
  "extractSocials": true,
  "verifyEmailDomains": true,
  "includePageDetails": false,
  "maxSites": 0,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `socials` (type: `string`):

No description

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "hawksheadrelish.com",
        "puccinibomboni.com",
        "fairmountbagel.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/website-contact-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "hawksheadrelish.com",
        "puccinibomboni.com",
        "fairmountbagel.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("gazidev/website-contact-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "hawksheadrelish.com",
    "puccinibomboni.com",
    "fairmountbagel.com"
  ]
}' |
apify call gazidev/website-contact-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/website-contact-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kxbtNOyghxgwCICba/builds/N4zQKXiEVq6pVcrdF/openapi.json
