# Website Contact Extractor (`cynix_dev/website-contact-scraper`) Actor

Crawl seed websites and extract contact details — emails, phone numbers, and social profiles (Instagram, X, LinkedIn, Facebook). Performs a light same-host crawl up to a per-domain page budget. Ideal for lead enrichment and outreach list building.

- **URL**: https://apify.com/cynix\_dev/website-contact-scraper.md
- **Developed by:** [Cynix Dev](https://apify.com/cynix_dev) (community)
- **Categories:** Lead generation, Automation, AI
- **Stats:** 2 total users, 1 monthly users, 58.3% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.25 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website Contact Extractor

Crawl seed websites and extract **contact details** — emails, phone numbers and social profiles (Instagram, X, LinkedIn, Facebook). Performs a light same-host crawl up to a per-domain page budget. Built for lead enrichment and outreach-list building.

### What it does

Given a list of starting URLs, this Actor crawls each domain (staying on the same host) up to a page budget, scans pages for contact signals, and returns one record per domain with every email, phone number and social profile it found.

It's the enrichment step after you have a list of company sites: feed it the homepages, get back the contacts.

### Features

- **Seed-URL input** — homepages or specific pages like `/contact`.
- **Light same-host crawl** — follows internal links up to `maxPages` per seed.
- **Contact extraction** — emails, phone numbers, and Instagram/X/LinkedIn/Facebook profiles.
- **Optional proxy** — enabled only if a site blocks datacenter IPs.
- **Clean typed fields** — base URL, host, emails, phones, and each social handle.

### What people use it for

- Lead enrichment — turn a list of company sites into contact records.
- Outreach list building — collect emails and socials for sales.
- Partnership sourcing — find who to talk to at a target company.
- Directory and index building.
- Journalist and PR contact discovery.

### Getting good coverage

Contact details often live on `/contact`, `/about` or the footer rather than the homepage. Pointing a seed directly at `https://example.com/contact` costs less crawl budget than the homepage, but a homepage seed with a reasonable `maxPages` will usually reach those pages too.

#### Data hygiene

Expect noise: some "emails" are `noreply@` or role addresses, and phone formats vary by country. Design your downstream step to:

- Drop `noreply`/role addresses if you want humans.
- Normalise phone numbers to E.164 before dialling.
- Dedupe by `host` across seeds that belong to the same company.

### Input

`startUrls` is required. `maxPages` sets the crawl budget per seed domain.

| Field | Type | Default | What it does |
| --- | --- | --- | --- |
| `startUrls` **(required)** | array | `["https://example.com"]` | Seed URLs to crawl, e.g. \["https://example.com", "https://another.com/contact"]. |
| `maxPages` | integer | `30` | Crawl budget per seed URL (same-host links only). Range 1–500. |
| `proxyConfiguration` | object | see below | Some sites block datacenter IPs. Enable Apify proxy if needed. |

#### Input example

```json
{
  "startUrls": [
    "https://example.com"
  ],
  "maxPages": 30
}
```

### Output

One record per seed domain: base URL, host, emails, phones and the social profiles found.

Every dataset record contains: `baseUrl`, `host`, `emails`, `phones`, `instagram`, `twitter`, `linkedin`, `facebook`, `pagesCrawled`, `fetchedAt`.

Export the dataset as JSON, CSV, Excel, XML or JSONL from the Console, or pull it programmatically through the Apify API and any of the official clients.

### How to use it

1. Click **Try for free** (or **Start** if you already have an Apify account).
2. Fill in the input fields described above — the defaults already produce a working run.
3. Press **Start** and watch the log; results stream into the dataset as they are found.
4. When the run finishes, open the **Output/Storage** tab and export as JSON, CSV or Excel.

Runs can be scheduled (hourly, daily, weekly) and wired into Slack, Google Sheets, Zapier, Make, webhooks or your own backend through Apify integrations. Everything the Console does is also available over the [Apify API](https://docs.apify.com/api/v2).

### Proxy configuration

This Actor accepts a standard Apify **proxy configuration** object. Residential proxy is the default because the target site rate-limits datacenter IP ranges; you can select a specific exit country or supply your own proxy URLs.

```json
{
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

### Pricing

This Actor is billed on Apify's **pay-per-event** model: a small charge when a run starts, plus a charge for each result written to the dataset. You only pay for records you actually receive — a run that finds nothing costs only the start event. Current rates are always shown on the **Pricing** tab of this page, and the run log prints your usage as it goes.

Free-plan credits from Apify cover a large amount of light usage, so you can evaluate the Actor before committing to anything.

### FAQ

#### Do I need a proxy?

Usually not — most sites serve their own pages to any IP. Enable `proxyConfiguration` only if a specific site blocks datacenter IPs.

#### How deep does it crawl?

`maxPages` caps pages per seed, and it stays on the same host. It's a light contact-discovery crawl, not a full-site archive.

#### Can it find emails behind a form?

No — only emails present in page text or `mailto:` links are captured. Emails gated behind a contact form aren't extractable.

#### Will it respect robots.txt?

It's a focused contact crawler; if you need strict robots compliance for a given site, scope `startUrls` and `maxPages` accordingly.

### Other Actors by cynix\_dev

| Actor | What it does |
| --- | --- |
| [Company Tech Stack & Hiring Intelligence](https://apify.com/cynix_dev/company-tech-intel) | Detect technologies from company websites, extract tech requirements from job postings, discover competitors. |
| [Job Postings — ATS Boards Extractor](https://apify.com/cynix_dev/job-postings-ats) | Pull live job postings straight from companies' public applicant-tracking boards — Greenhouse, Lever, Ashby, SmartRecruiters, and … |
| [SEC EDGAR Filings Extractor](https://apify.com/cynix_dev/sec-edgar-filings) | Search and extract SEC EDGAR filings: full-text search across all filings or company filing histories by CIK. |
| [WHOIS & DNS Enrichment](https://apify.com/cynix_dev/whois-enrichment) | Enrich domains with structured WHOIS data (registrar, registration/expiration dates, status, nameservers) and optional DNS … |
| [News & Press Release Monitor](https://apify.com/cynix_dev/news-press-monitor) | Watch company newsrooms, blogs, and press pages and get one clean record per article — with new-item detection between runs, so a … |

### Legal and responsible use

This Actor collects only publicly available information. You are responsible for how you use the data, including compliance with the target site's Terms of Service, robots directives, copyright, and data protection law such as GDPR and CCPA. Do not use it to gather personal data without a lawful basis.

### Support and feedback

Found a bug, hit a site change, or need an extra field? Open a ticket on the **Issues** tab of this Actor — issues are read and fixed. Feature requests and custom-scraper enquiries are welcome through the same channel.

# Actor input Schema

## `startUrls` (type: `array`):

Seed URLs to crawl, e.g. \["https://example.com", "https://another.com/contact"].

## `maxPages` (type: `integer`):

Crawl budget per seed URL (same-host links only).

## `proxyConfiguration` (type: `object`):

Some sites block datacenter IPs. Enable Apify proxy if needed.

## Actor input object example

```json
{
  "startUrls": [
    "https://example.com"
  ],
  "maxPages": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `baseUrl` (type: `string`):

Page URL where contacts were found.

## `host` (type: `string`):

Website host.

## `emails` (type: `string`):

Semicolon-separated emails.

## `phones` (type: `string`):

Semicolon-separated phone numbers.

## `instagram` (type: `string`):

Instagram profile URL.

## `twitter` (type: `string`):

X/Twitter profile URL.

## `linkedin` (type: `string`):

LinkedIn profile/company URL.

## `facebook` (type: `string`):

Facebook profile URL.

## `pagesCrawled` (type: `string`):

Pages crawled for this host.

## `fetchedAt` (type: `string`):

ISO timestamp.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("cynix_dev/website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("cynix_dev/website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call cynix_dev/website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cynix_dev/website-contact-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NzlCxUgSzMGu4MtVQ/builds/MNYuLAf5Va3VDr7jK/openapi.json
