# Website Contact Enricher — Emails, Phones & Source Evidence (`samstorm/website-contact-enricher`) Actor

Enrich business website lists with public emails, phones, social links and contact-page evidence. Preserve input IDs, distinguish other-location contacts, and report blocked or incomplete scans. No external API key required.

- **URL**: https://apify.com/samstorm/website-contact-enricher.md
- **Developed by:** [Sam Kleespies](https://apify.com/samstorm) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 website scanneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Enricher

Turn a list of business websites into **public emails, normalized phone numbers, linked social profiles and source evidence**. Keep your CRM row IDs, inspect the context around each contact, and see exactly which websites were blocked or only partly scanned.

Use this after collecting company websites from directories, event exhibitor lists or your own CRM. This Actor enriches supplied websites; it does not discover companies, find private employee emails, send messages or verify email delivery. **Public beta: test a small batch against your workflow before scaling.**

Researching an event? [Map Your Show Exhibitor Scraper](https://apify.com/samstorm/map-your-show-exhibitor-scraper#enrich-selected-exhibitor-websites) provides an optional first step and a website-handoff recipe. Each Actor is priced separately.

### Quick start

1. Paste your list into **Website URLs**, one per entry. A homepage, domain or specific branch page works.
2. Optionally add a two-letter **Phone country hint**, such as `US` or `GB`, to normalize national telephone numbers.
3. Keep the three-page default, set your maximum run spend, and select **Start**.
4. Open **Contacts and coverage**, then **Sources and uncertainty** when you need to inspect a value. Export JSON for nested evidence or CSV for the flat email/phone columns.

The prefilled example uses two public business contact pages. Website content and access can change between runs.

```json
{
  "urls": [
    "https://www.pikeplacechowder.com/pikeplacemarketlocation",
    "https://www.pololu.com/contact"
  ],
  "country": "US",
  "maxPages": 3,
  "concurrency": 2,
  "runTimeoutSecs": 60
}
```

### What makes the output useful?

- **Evidence for every contact:** page URL, extraction method, observation time and nearby source text.
- **One row per input:** preserve duplicates and IDs rather than silently dropping unsuccessful websites. Completion order can differ; `inputIndex` restores input order.
- **Conservative suggestions:** optional primary email and phone fields can stay blank when contacts are ambiguous, special-purpose, conflicted or not supported by the supplied branch hint.
- **Purpose hints:** distinguish general, sales, support, reservation, press, recruitment, privacy and executive-escalation contexts. These are heuristics, not verified business classifications.
- **Phone normalization:** equivalent supported formats are consolidated; extensions and link/display disagreements remain available.
- **Branch evidence:** contacts with explicit structured evidence for other locations are separated into `otherLocationEmails` and `otherLocationPhones`. This does not classify every unstructured location block.
- **No external API key:** no funded email database, SMTP checker, browser-rendering service or residential proxy is required.

### Pricing

**$0.02 per run start + $0.005 per website where at least one usable public HTML page is inspected ($5 per 1,000). Apify platform usage is included.**

A readable website counts once regardless of the number of contacts or pages within your selected limit. Readable websites with no observed contacts and partially readable websites also count. Blocked, failed, invalid and unscanned diagnostic rows have **no website-result fee**. The start fee still applies, including runs with no readable websites.

| Usable websites | Maximum event charge for that run |
| --- | --- |
| 2 | $0.03 |
| 10 | $0.07 |
| 50 | $0.27 |
| 100 | $0.52 |

These examples assume one run and no pricing discount. The live Pricing tab is authoritative. Use a maximum cost of at least **$0.025** for the start and one website. The Actor reserves spending capacity before fetching a website, stops starting work when the allowance is exhausted and saves explicit unscanned diagnostics. Duplicate input entries can each incur a website fee.

### Advanced input: preserve CRM IDs or target a branch

Use **Advanced website records** instead of Website URLs. Clear the URL list before supplying records; sending both lists is rejected to prevent accidental double processing.

```json
{
  "targets": [
    {
      "website": "https://www.pikeplacechowder.com/pikeplacemarketlocation",
      "inputId": "crm-row-42",
      "businessName": "Pike Place Chowder",
      "country": "US",
      "locationHint": "Pike Place Market"
    }
  ],
  "maxPages": 3
}
```

`inputId` and `businessName` are your supplied values, not independently verified identities. `country` on a record overrides the shared country hint. A location hint is a city or branch label, matched against available source context. Supply the actual branch-page URL when possible. A company-wide or alternate-domain contact may remain in the secondary list without a primary suggestion.

| Input | Limits / default |
| --- | --- |
| `urls` or `targets` | 1–100 entries; choose one list |
| `country` | Optional uppercase ISO country code |
| `maxPages` | 1–5 attempted content pages per website; default 3 |
| `concurrency` | 1–2 websites; default 2 |
| `runTimeoutSecs` | 10–600 seconds of work; default 300 |

The page limit counts attempted content pages. Robots checks and redirects are tracked separately. Each website has a 50-second collection budget. Parsing and saving can extend elapsed time; set the external run timeout at least 60 seconds beyond the work budget. Batches of 100 slow sites may need to be split. Per-page and per-site byte limits bound unusually large responses.

### From an exhibitor shortlist to contact evidence

Start with an exported exhibitor shortlist. Select extracted `matched` profiles, or `not_requested` when you intentionally used no filters; do not silently treat unknown categories as matches. Review and deduplicate identical website/name pairs while keeping original profile IDs. Different branch URLs or names should remain separate.

Clear Website URLs, then map `website` to `targets[].website`, `profileId` to `targets[].inputId`, and `name` to `targets[].businessName`. Keep each batch at 100 or fewer records. Join output by `inputId`, because completion order can differ. If you grouped duplicates, preserve the one-to-many mapping outside this Actor. Do not infer a phone country from the event's location.

An owner test on September 30, 2026 UTC reused a completed exhibitor export and sent five matched websites through this Actor. All five IDs returned: four websites yielded usable HTML, two had observed emails and three had phones. One website redirected to another domain and received a free diagnostic row. This is a small technical handoff test, not evidence of paid customer demand or guaranteed email coverage. Inspect any alternate domain before supplying a replacement URL.

The contact stage would cost $0.04 at standard event prices for one start and four readable websites; this is a customer-price illustration, not the developer's own compute bill. Readable websites still count when no contact is observed. The earlier exhibitor stage has separate charges.

### Read results correctly

| Field | Meaning |
| --- | --- |
| `emailValues`, `phoneValues` | Flat, deduplicated contact lists for CSV workflows |
| `emails`, `phones` | Contact records with per-value evidence and uncertainty |
| `primaryEmail`, `primaryPhone` | Optional suggestions; not verified contact destinations |
| `primaryEmailStatus` | Suggested from source evidence, ambiguous, or not established |
| `emailSourceUrls` | Source URLs behind the email list |
| `socials` | Profiles linked from the site; ownership unverified |
| `contactForms` | Detected forms; never submitted |
| `otherLocationEmails`, `otherLocationPhones` | Explicit other-location structured contacts, where detectable |
| `conflicts` | Observed telephone link/display disagreements |
| `pages`, `robots`, `metrics` | Actual page coverage, access checks and work metrics |
| `billing` | Whether the row represents a scanned-website event or a free diagnostic |

`status` is one of:

- `complete_within_budget`: attempted content pages were usable. This **does not mean the whole website was crawled**; inspect `pendingPages` and `stopReason`.
- `partial`: at least one usable page, with a failed page or collection limit.
- `blocked`: access policy, challenge or robots checks prevented collection.
- `failed`: no usable supported HTML was obtained.
- `invalid_input`: unsafe or unsupported website URL.
- `not_scanned`: work or spending limit prevented starting this input.

A blank email list means **not observed on the inspected pages**, not that the business has no email. `deliverability: "not_checked"` and `reachability: "not_checked"` are intentional. Unusual source addresses are preserved with warnings, not silently repaired. A blank primary can coexist with useful secondary contacts.

The **Run coverage and work metrics** output link opens `RUN_SUMMARY`, including statuses, readable website count and unscanned input indexes. To continue, start a new run with only the unscanned/unfinished records after reviewing prior output. Resurrect is unsupported to avoid repeat output and charges. Completed rows are saved incrementally if a later operation fails.

### Example output

Abbreviated owner-validation result observed September 30, 2026 UTC. Only one of the returned email records is shown. This is a source-observation example, not a deliverability test.

```json
{
  "inputId": "ce-033",
  "website": "https://www.pololu.com/",
  "status": "complete_within_budget",
  "primaryEmail": "inbox@pololu.com",
  "primaryEmailStatus": "suggested_from_source_evidence",
  "emails": [
    {
      "value": "inbox@pololu.com",
      "domainRelationship": "same_domain",
      "evidence": [
        {
          "sourceUrl": "https://www.pololu.com/contact",
          "method": "mailto",
          "raw": "inbox@pololu.com",
          "context": "General questions and comments \u2013 inbox@pololu.com",
          "observedAt": "2026-09-30T01:24:16.775Z",
          "purpose": "general",
          "scope": "unspecified"
        }
      ],
      "warnings": [],
      "deliverability": "not_checked",
      "purposes": [
        "general"
      ],
      "scopes": [
        "unspecified"
      ]
    }
  ]
}
```

### Scope and limitations

This is a static HTTP crawler. It reads supported public HTML, mail/telephone links, visible text, published email obfuscation and selected organization structured data. It follows observed same-host contact/about/location links. It does not execute page JavaScript, guess hidden contact paths, bypass access challenges, submit forms or use a private contact database.

Robots restrictions and unknown robots responses stop collection; sites with crawl delays beyond the budget are reported explicitly. Private-network URLs, credential-bearing URLs, custom ports and unrelated-domain redirects are refused. Use a company's final canonical URL if it moved to a different domain.

Source context, contact purpose and branch association are heuristic. Different websites require different methods; this beta can yield fewer reachable sites than a less restrictive crawler. Phone-format validation is not proof that a number belongs to the business. Use source evidence to decide whether a contact is appropriate for your purpose.

### API, automation and AI agents

The **API** tab provides generated Python, JavaScript and HTTP examples for `samstorm/website-contact-enricher`. Send the same input JSON, wait for the run to finish, then read its dataset. Check row statuses and `RUN_SUMMARY`, rather than treating a successful run as success on every website.

You can use Apify's native task, schedule and integration features after confirming your input. For an AI assistant, open this Actor's native MCP setup and select it in Apify's hosted MCP server. Apify authentication and usage charges still apply; no separate data-provider key is needed. An external authenticated MCP client has not been validated for this beta.

### Support

Open an issue with the run ID, input website, expected field and a public source URL showing the discrepancy. Include the country or branch hint when relevant. Do not include passwords or API tokens. Reproducible failures and misleading output take priority over expanding the crawl scope.

### Changelog

**0.2.0 — public beta:** simpler URL input, CRM record alternative, source/evidence/diagnostic views, conservative primary suggestions, explicit branch separation, normalized phones, bounded billing and incremental output. Private tests are owner validation, not evidence of paid customer demand or guaranteed accuracy.

# Actor input Schema

## `urls` (type: `array`):

Paste public business homepages or branch URLs. Domains without https:// are accepted. Up to 100 per run. Each entry remains a separate record.

## `country` (type: `string`):

Uppercase two-letter ISO code, for example US or GB. Helps parse national phone numbers. Leave empty for international numbers or set country per advanced record.

## `maxPages` (type: `integer`):

Counts attempted content pages. Robots checks and redirects are tracked separately. Reaching this cap does not mean the whole site was inspected.

## `targets` (type: `array`):

Optional alternative to Website URLs. Objects: website, inputId, businessName, country, locationHint. Do not supply both lists.

## `concurrency` (type: `integer`):

At most two targets at a time, with same-site pacing.

## `runTimeoutSecs` (type: `integer`):

Stops starting new work at this limit. Finishing/parsing a page can extend elapsed time. Unscanned rows are explicit; retry only those inputs. Keep the external run timeout at least 60 seconds longer.

## Actor input object example

```json
{
  "urls": [
    "https://www.pikeplacechowder.com/pikeplacemarketlocation",
    "https://www.pololu.com/contact"
  ],
  "country": "US",
  "maxPages": 3,
  "concurrency": 2,
  "runTimeoutSecs": 300
}
```

# Actor output Schema

## `contacts` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.pikeplacechowder.com/pikeplacemarketlocation",
        "https://www.pololu.com/contact"
    ],
    "country": "US"
};

// Run the Actor and wait for it to finish
const run = await client.actor("samstorm/website-contact-enricher").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.pikeplacechowder.com/pikeplacemarketlocation",
        "https://www.pololu.com/contact",
    ],
    "country": "US",
}

# Run the Actor and wait for it to finish
run = client.actor("samstorm/website-contact-enricher").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.pikeplacechowder.com/pikeplacemarketlocation",
    "https://www.pololu.com/contact"
  ],
  "country": "US"
}' |
apify call samstorm/website-contact-enricher --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,samstorm/website-contact-enricher"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pVOwgBHxVDL6VRE0z/builds/GdXYVW2LNPZ1Xb21d/openapi.json
