# Website Contact Scraper (`wildorigins/website-contact-scraper`) Actor

🏷️ From $0.60 / 1K | Turn company websites into contact records. Finds emails split into personal and role, phone numbers, social profiles across 17 platforms, postal address and named people. Renders JavaScript sites at no extra cost. One charge per domain, and domains with nothing found are free.

- **URL**: https://apify.com/wildorigins/website-contact-scraper.md
- **Developed by:** [Wild Origins](https://apify.com/wildorigins) (community)
- **Categories:** Lead generation, Automation
- **Stats:** 4 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scraper

Company websites in, one contact record per domain out: emails, phone numbers, social profiles, postal address, opening hours and named people.

### 🔍 What does Website Contact Scraper do?

It reads a company's own website and pulls out the contact details published on it, then returns one record per registrable domain so www and subdomains collapse into a single row.

One charge per domain that produced contact details, whatever the size of the site. **Domains where nothing is found are not charged.**

#### Exactly what it extracts

| | From | Notes |
|---|---|---|
| **Email addresses** | `mailto:` links and page text | Split three ways, see below |
| **Phone numbers** | `tel:` links and page text | Each says which it came from |
| **Social profiles** | Links anywhere on the page | 17 platforms, listed below |
| **Postal address** | schema.org markup only | Street, city, region, postcode, country |
| **Opening hours** | schema.org markup only | As published |
| **Named people** | schema.org `Person` markup only | Name, job title, email where given |

The 17 platforms: LinkedIn, X, Facebook, Instagram, YouTube, GitHub, TikTok,
Pinterest, Threads, Bluesky, Mastodon, Crunchbase, Reddit, Discord, Telegram,
WhatsApp and Snapchat.

### 📊 What data can I extract from a website?

One record per domain:

| Field | What it is |
|---|---|
| `domain` | Registrable domain, so www and subdomains collapse to one record |
| `hasContactData` | Whether anything was found. False is never charged |
| `emails` | Each with `type`, `isFreeProvider`, `onSiteDomain`, `foundOn` |
| `emailCount` | With `personalEmailCount`, `roleEmailCount`, `unclassifiedEmailCount` |
| `phones` | Each with `source`, `tel-link` or `text`, and `foundOn` |
| `socials` | Profile URLs keyed by platform |
| `people` | Name, job title, email where published |
| `address`, `hours` | From schema.org, with the page each came from |
| `pagesCrawled`, `pagesVisited` | Exactly what was read to produce the record |
| `queryUrl`, `startUrl` | The URL asked for, and the one the scan began from after redirects |
| `phoneCount` | How many phone numbers were found |
| `checkedAt` | When the scan ran, ISO 8601 |
| `error` | Set only when no page could be read at all |

A domain that returns nothing still produces a record with `hasContactData:
false`, unless you switch that off. That is a real finding, not a failure: it
tells you the company publishes no contact details.

### 💡 Why scrape website contact details?

**B2B outreach.** Build a contact list from what businesses publish themselves, with role and personal addresses kept apart so your first line lands right.

**Lead enrichment.** Take a list of domains you already have and attach phones, socials and addresses to each one.

**Local business research.** Collect opening hours and postal addresses from schema.org markup across a sector or region.

**List hygiene.** See which domains publish nothing at all before you spend anything on them.

### 🚀 How do I use Website Contact Scraper?

1. Click **Try for free**.
2. Put a website into `url`, or up to 500 into `urls`.
3. Set `maxPagesPerDomain`, and turn on `deepScan` if the sites have sparse menus.
4. Click **Start** and wait for the run to finish.
5. Download the results as JSON, CSV or Excel, or pull them from the API.

### ⬇️ Input

```json
{
  "url": "monzo.com",
  "maxPagesPerDomain": 5
}
```

| Field | Type | Default | What it does |
|---|---|---|---|
| `url` | string | `monzo.com` | A single website, protocol optional |
| `urls` | array | | Up to 500 websites, charged per domain |
| `maxPagesPerDomain` | integer | `5` | Pages to crawl per domain, 1 to 20 |
| `deepScan` | boolean | `false` | Also probe standard contact paths |
| `includeFreeProviders` | boolean | `true` | Keep gmail.com and similar addresses |
| `onlyWithContacts` | boolean | `false` | Drop empty domains from the output |
| `timeoutSecs` | integer | `20` | Per page timeout |

### ⬆️ Output

#### Table view

Results arrive as a Sites table you can sort and filter in the Console, with the domain, whether contact data was found, the addresses and phones, how many social platforms were seen, how many people were named and the page count lined up per domain.

#### JSON

A typical row:

```json
{
  "domain": "monzo.com",
  "hasContactData": true,
  "emails": [
    { "email": "help@monzo.com", "local": "help", "domain": "monzo.com", "type": "role" }
  ],
  "phones": [
    { "phone": "+44 20 7946 0000", "source": "tel-link", "foundOn": "https://monzo.com/help" }
  ],
  "socialPlatformCount": 5,
  "peopleCount": 0,
  "pagesCrawled": 5,
  "emailCount": 1
}
```

Download it from the run as JSON, CSV or Excel, or read it straight from the API.

### Three address types, not two

Most tools split addresses into personal and role. That forces a guess on every
single word local part, and the guess is wrong often enough to matter.

| Type | Meaning |
|---|---|
| `role` | A known department word: `info`, `sales`, `support`, `careers` |
| `personal` | Name shaped, two parts around a separator: `jane.smith`, `j.smith` |
| `unclassified` | A single word that could be either: `greg`, `solar`, `heat` |

`solar@`, `heat@` and `greg@` are the same shape. Two are departments and one is
a person, and nothing in the address says which. Calling a department mailbox a
named human is the expensive direction of that mistake, because the first line
of your outreach lands wrong. So those come back `unclassified` rather than
assigned.

Every address also carries `isFreeProvider`, and `onSiteDomain`, which is false
when the address sits on a different domain to the site. That is how you tell a
company's own address from its agency's.

### How it reads a site

The homepage first, then the pages most likely to carry contact details, ranked:
contact, imprint, about, team, then support and legal. A shallow `/contact`
beats a deep `/blog/2019/contact-us-update`. Same registrable domain only, so a
link to a partner's site is never followed.

`deepScan` also tries the standard contact paths directly, including the imprint
and legal notice pages European sites are required to publish. Useful when a
site's menu is sparse or built in JavaScript.

### What it does not do

It does not run a browser. Every page is read over plain HTTP, which is what
keeps a run fast and cheap.

A browser fallback used to render pages whose HTML carried no contacts. It was
removed because it was not earning its place: it fired on none of the last ten
runs, and on the client rendered sites it existed for, the contact details are
absent from the HTML and from the page's own JSON payloads alike, because those
sites publish a contact form rather than an address. A browser cannot conjure an
address that is not there.

So a site that publishes contact details only after its scripts run will come
back empty here. It also comes back **free**, because a domain that produced
nothing is never charged.

### What it deliberately does not do

Each of these is a refusal to guess, because in an outreach list a wrong answer
costs more than a missing one.

- **Names are never inferred from headings.** Only schema.org `Person` markup is
  read. Guessing from `<h2>` returns product names and page furniture presented
  as staff.
- **Bare digit runs are never read as phone numbers.** A price, an order number
  and a company registration number are indistinguishable from a phone number to
  a loose pattern.
- **Addresses are never inferred from text.** Without schema.org markup, the
  address is reported absent.

It also does not: verify mailboxes by SMTP, use a paid enrichment database, read
login gated pages, or follow links off the domain. It finds what a company chose
to publish on its own website, and nothing beyond that.

### Run timeout

A run stops starting new work shortly before its own time limit and finishes cleanly with whatever it has, naming the sites it did not reach. The Actor's default is 3600 seconds, which is enough for the largest input it accepts. If you set a shorter limit in your own run settings, integration or API call, expect fewer results and a note in the run's status message saying so. You pay per delivered result rather than per minute, so a generous timeout costs you nothing.

### ⏱️ How long does a run take?

Measured on real runs, so you can tell a normal run from one that has stalled.

| Input size | Typical run time |
|---|---|
| 1 site | 3 to 8 seconds |
| 50 sites | about 5 minutes |

The first few seconds of any run are the container starting rather than the work. A run is never silently stuck: progress is logged as it goes, and if it runs out of time it stops early, keeps everything collected so far and says in the status message what was left.

### 💰 How much does it cost?

| Event | Price |
|---|---|
| Website scanned | $0.001 |
| Actor start | $0.00005 per run |

Charged once per domain that produced contact details. Not per page and not per
address found. A run over 100 sites
with a 60% hit rate costs 60 at $0.001 plus one run start, **$0.06005**, not
$0.10005.

Paid Apify plans pay less per website: **$0.00085** on Bronze, **$0.0007** on
Silver, **$0.0006** on Gold, **$0.0005** on Platinum and **$0.0004** on
Diamond, which is 40 percent of the list price. The Apify listing always shows
the current rates.

No API keys, no accounts, no proxies.

### 🔌 Integrations

Send results straight to Google Sheets, Slack, Airtable, Zapier, Make or your own webhook using [Apify integrations](https://docs.apify.com/platform/integrations). You can also trigger a run whenever something happens in another tool.

### 🔗 Using Website Contact Scraper with the Apify API

```bash
curl -X POST "https://api.apify.com/v2/acts/spookyweb~website-contact-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"url": "monzo.com", "maxPagesPerDomain": 5}'
```

Or with the Apify client:

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('spookyweb/website-contact-scraper').call({
  url: 'monzo.com',
  maxPagesPerDomain: 5,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

Full detail is in the [Apify API reference](https://docs.apify.com/api/v2), and every run is also callable from the [Python](https://docs.apify.com/api/client/python) and [JavaScript](https://docs.apify.com/api/client/js) clients.

### ❓ FAQ

#### Which pages does it read?

The homepage first, then the pages most likely to carry contact details, ranked: contact, imprint, about, team, then support and legal. A shallow `/contact` beats a deep `/blog/2019/contact-us-update`. It stays on the same registrable domain, so a link to a partner's site is never followed. `maxPagesPerDomain` caps how many are read, and `deepScan` also probes the standard contact paths directly.

#### Does it render JavaScript?

No. Every page is read over plain HTTP. A browser fallback was removed on 31 August 2026 after measurement showed it firing on none of the last ten runs and recovering nothing on the sites it was written for, which publish a contact form rather than an address. A site that only reveals contact details after its scripts run returns an empty record here, and an empty record is never charged for.

#### Are free provider addresses like gmail included?

Yes by default, and every address carries `isFreeProvider` so you can filter them out yourself. Set `includeFreeProviders` to false to drop them during the run instead. Small businesses often publish a Gmail address as their only contact, so throwing them away by default would lose real data.

#### What are the three address types?

`role` is a known department word such as info, sales, support or careers. `personal` is name shaped, two parts around a separator such as `jane.smith`. `unclassified` is a single word that could be either, such as `greg`, `solar` or `heat`. Those stay unclassified rather than being guessed at, because calling a department mailbox a named human is the expensive direction of the mistake.

#### Are domains with nothing found still charged?

No. The charge only happens when a domain produced contact details. A run over 100 sites with a 60% hit rate costs $0.06 rather than $0.10. The empty domains still appear in the dataset with `hasContactData: false`, unless you set `onlyWithContacts` to drop them.

#### Can it find named people?

Only where the site publishes schema.org `Person` markup, in which case you get the name, job title and email where given. Names are never inferred from headings, because guessing from an `<h2>` returns product names and page furniture presented as staff.

### ⚖️ Is it legal to scrape website contact details?

This reads contact details a business has published on its own website so customers can reach it, which is business contact information. It reads only public pages and never anything behind a login.

Where a detail identifies a named person, UK GDPR applies and you are the data controller for what you do next. You need a lawful basis, typically legitimate interest for B2B contact, and you must honour opt outs. Apify's [ethical scraping guide](https://blog.apify.com/is-web-scraping-legal/) covers the wider picture.

### 👍 Your feedback

Found a bug, or want a field that is not here yet? Open an issue on the Actor's Issues tab. Requests that make the data more useful get built, and problems get fixed quickly.

### 🔎 You might also like

| Actor | What it does |
|---|---|
| [Company Email Finder](https://apify.com/spookyweb/company-email-finder) | Published company addresses, the naming pattern behind them and an MX check |
| [Wayback Machine Scraper](https://apify.com/spookyweb/wayback-machine-scraper) | Every archived capture of a URL, what changed between them, and the original bytes |
| [UK Food Hygiene Ratings](https://apify.com/spookyweb/uk-food-hygiene-ratings) | Food hygiene ratings for every UK food business, straight from the FSA |

# Actor input Schema

## `url` (type: `string`):

A single company website to scan. The protocol is optional, so acme.com and https://acme.com are the same request.

## `urls` (type: `array`):

Scan up to 500 company websites in one run. Takes priority over the single website field when both are given. Each domain is charged once, whatever the size of the site.

## `maxPagesPerDomain` (type: `integer`):

How many pages to read on each site. The homepage is always read, then the pages most likely to carry contact details, ranked ahead of the rest: contact, imprint, about, team, then support and legal.

## `deepScan` (type: `boolean`):

Also try the usual contact URLs directly even when the site does not link to them, including the imprint and legal notice pages that European sites are required to publish. Slower, and worth it on sites with a sparse or JavaScript driven menu.

## `includeFreeProviders` (type: `boolean`):

Keep addresses on free consumer providers such as gmail.com and outlook.com. Useful for sole traders and small businesses that publish a personal address, and noise if you only want corporate domains.

## `onlyWithContacts` (type: `boolean`):

Drop domains where nothing was found instead of returning an empty record for them. Those domains are never charged either way, so this only affects whether they appear in the output.

## `timeoutSecs` (type: `integer`):

How long to wait for a single page before moving on. Raise it for slow sites, lower it to get through a large list faster.

## `concurrency` (type: `integer`):

How many websites to read at the same time, from 1 to 20. Raising it gets through a long list faster. Four suits most lists. Raise it for lists in the hundreds.

## Actor input object example

```json
{
  "url": "monzo.com",
  "maxPagesPerDomain": 5,
  "deepScan": false,
  "includeFreeProviders": true,
  "onlyWithContacts": false,
  "timeoutSecs": 20,
  "concurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

One row per item: websites with the contact details found.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "url": "monzo.com"
};

// Run the Actor and wait for it to finish
const run = await client.actor("wildorigins/website-contact-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "url": "monzo.com" }

# Run the Actor and wait for it to finish
run = client.actor("wildorigins/website-contact-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "url": "monzo.com"
}' |
apify call wildorigins/website-contact-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,wildorigins/website-contact-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/VfKrgOcFVNTClVHs8/builds/ALw9BFDZodt07qil4/openapi.json
