# Website Contact Scout — Emails, Phones & Sources (`iwins/website-contact-scout`) Actor

Find published website emails, phone links and social profiles with source-page evidence. Check linked contact pages, deduplicate results and distinguish blocked sites from successful scans. Static HTML, no AI API required.

- **URL**: https://apify.com/iwins/website-contact-scout.md
- **Developed by:** [Ahmed Firas](https://apify.com/iwins) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 website checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contact Scout

Turn a list of business websites into a compact, source-backed contact report. Find published email addresses, clickable phone numbers and linked social profiles, then see exactly which page supplied each result.

Use it to enrich website URLs from your own business directory, an authorized lead export or a CRM cleanup workflow. It does not search Google Maps itself, send messages, guess email addresses or test mailboxes.

### What makes this useful

- One consolidated result per website, rather than a separate row for every page.
- Source URL and extraction method for every email, phone link and social profile.
- Follows linked contact, about and imprint pages within the page limit.
- Labels common role inboxes such as `sales@` and `support@`; other addresses remain unclassified.
- Deduplicates repeated website domains and contact values.
- Distinguishes a successful scan with no contacts from a blocked, unavailable or disallowed website.
- Reports partial coverage when secondary pages fail. No invented contacts or confidence scores.

### Run it

1. Add up to **20 public HTTPS website URLs** in Website URLs.
2. Choose **1–4 pages per website**. The starting page counts toward the limit. Only discovered contact/about/imprint links are considered for additional pages.
3. Start the run. The example checks one public Python.org help page; it is a demonstration, not a customer or sales lead.
4. Open **Output → Website contacts** and export JSON, CSV or Excel using Apify's export controls.
5. Open **Evidence and failures → REPORT** for failed and unattempted websites. `SITE-001`, etc. hold the saved per-website evidence.

```json
{
  "startUrls": [{"url":"https://www.python.org/about/help/"}],
  "maxPagesPerSite": 1
}
```

### Output

| Field | Meaning |
| --- | --- |
| `website` | Supplied website after query/fragment removal. |
| `status` | `contacts_found` or `no_contacts_found` for a successfully checked page. |
| `emails` | Deduplicated published addresses. No mailbox/deliverability verification. |
| `phones` | Values from explicit `tel:` links, not guesses from arbitrary numbers. |
| `socialProfiles` | Linked social URLs; their contents are not crawled or verified. |
| `evidence` | Value, kind, source URL, extraction method and optional inbox/platform label. |
| `pagesChecked` | Successfully read HTML pages. |
| `warnings` | Explicit failures of secondary pages. |
| `checkedAt` | UTC observation time. |

If every website fails, the run fails with **no chargeable dataset results**. Read `REPORT` to see why. If some websites succeed, their results remain available alongside a report of failures. A site that loads but publishes no detectable contacts still counts as a successfully checked website.

### Pricing

Launch pricing: **USD 0.005 per successfully checked website** ($5 per 1,000 websites), including up to the selected page limit and platform usage. One dataset item equals one website scan, **not one email or one page**. Blocked/unavailable starting pages are not appended to the dataset and do not trigger a result charge. There is no startup charge or AI API fee. Check the live Pricing tab before running.

The maximum run cost limits the number of successful website results. For example, a $0.01 budget permits two websites at this price. Set at least $0.005. If a final storage request has an uncertain outcome, inspect the dataset before starting another run. A run with existing results cannot be resurrected to avoid duplicate result charges. Starting a new run is a new scan and may be charged again.

### Coverage and boundaries

- **Static HTML only.** No JavaScript rendering, login, CAPTCHA solving, proxy rotation or anti-bot bypass. JS-only contact details may be missed.
- `robots.txt` is checked before page requests. Disallowed pages are skipped; unavailable or restricted robots files produce an explicit failure. Requests are paced at least one second apart per domain and honor supported crawl delays up to ten seconds.
- One bounded retry for HTTP 429/503. Longer retry instructions end that check rather than ignoring the site's requested delay.
- HTTPS only, with public IPv4 DNS. Local/private IPs, credential-bearing URLs and nonstandard ports are rejected. Sites reachable only by IPv6 are unsupported.
- Redirects are limited and remain on the same hostname or its `www` counterpart. Other domain redirects are reported for review.
- URLs lose query strings and fragments; signed/token-based or query-routed pages are unsupported.
- Response limit: 2 MiB per request and a 12-second response deadline. No guarantee that all contacts on a site are found.
- Common HTML entities and percent-encoded `mailto:` addresses are supported. JavaScript/Cloudflare obfuscation and image-based contact details are unsupported.
- Common placeholder addresses and image-file false positives are excluded. Phone values are not normalized to E.164. A social URL is a published link, not proof of profile ownership.

Only process websites you are authorized to access and use results appropriately. Published contact information does not establish consent to receive marketing. No emails are sent, and extracted addresses are not tested using SMTP. Raw webpage HTML is not stored; extracted results remain in your Apify storage under its retention/access settings.

### Support

Open an issue with an error code and a public or fictional reproduction. Do not put confidential lists, credentials or identity documents into public issues.

# Actor input Schema

## `startUrls` (type: `array`):

One to twenty public HTTPS website URLs. Repeated domains are checked once. Use only websites you are authorized to process.

## `maxPagesPerSite` (type: `integer`):

Starting page plus linked contact/about/imprint pages. Includes failed page attempts; robots.txt requests are additional. Static HTML only.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.python.org/about/help/"
    }
  ],
  "maxPagesPerSite": 1
}
```

# Actor output Schema

## `contacts` (type: `string`):

No description

## `reports` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.python.org/about/help/"
        }
    ],
    "maxPagesPerSite": 1
};

// Run the Actor and wait for it to finish
const run = await client.actor("iwins/website-contact-scout").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.python.org/about/help/" }],
    "maxPagesPerSite": 1,
}

# Run the Actor and wait for it to finish
run = client.actor("iwins/website-contact-scout").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.python.org/about/help/"
    }
  ],
  "maxPagesPerSite": 1
}' |
apify call iwins/website-contact-scout --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,iwins/website-contact-scout"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zuK4D1yP41y2nffX6/builds/9bYvUOq2tXY83HfYX/openapi.json
