# Website Contact Finder | Emails, Phones & Source Evidence (`peerless_columbine/website-contact-finder`) Actor

Extract public business emails, phones and social profiles from websites. Follow contact pages, retain source evidence, preserve duplicate inputs and failure records. Clear page limits and optional email checks.

- **URL**: https://apify.com/peerless\_columbine/website-contact-finder.md
- **Developed by:** [tingyou333 zhuang](https://apify.com/peerless_columbine) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 readable website records

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Contacts | Emails, Phones & Source URLs

Turn a list of company websites into records of their publicly listed email addresses, phone numbers and social links. Every accepted contact includes its source page, extraction method, fetch time and page hash, so you can check the context before importing it into a CRM.

Use this Actor to review company contact pages, refresh published support channels or audit a website list. It reads public HTML and follows contact, about and company links on the same site. It does not guess email addresses, search for people or infer relationships across websites. A published address or social link can belong to a third party; its presence does not establish company ownership.

### Start with a small scan

Paste this input into the Actor's JSON input editor and run it:

```json
{
  "urls": ["https://www.scrapingbee.com/"],
  "maxPagesPerSite": 3,
  "maxConcurrency": 1,
  "maxWebsitesConcurrency": 1,
  "verifyEmails": false,
  "useProxy": false
}
```

Open the default dataset to view or export website records. Keep `scanStatus` when importing results: a readable page with no contacts and a failed scan both have empty contact arrays, but mean different things. The run's key-value store contains `OUTPUT` for totals and limits, and `SOURCE_DIAGNOSTICS` for individual input failures and delivery details.

Three pages are a useful small starting point, not a promise of complete contact coverage. Increase `maxPagesPerSite` for deeper discovery or start directly at a public contact/legal page. Omitting this option uses 20 pages. Source content and access can change between runs.

### What a contact record looks like

This is an excerpt from an actual Hetzner scan on **2026-09-26 UTC**, version **0.1.3**, build `mGDQUJDwzRxITNXbg`, run `tqMJFBi3BnroHGP8s`. It preserves the recorded values and shows only one email, one phone and that email's complete evidence entry; other contacts and fields are omitted. Eighteen pages were readable and two failed, so the row is correctly marked `partial`. This older sample illustrates extraction and provenance; it is not a fresh scan or a claim about every website.

```json
{
  "websiteUrl": "https://www.hetzner.com/de/",
  "contactPageUrl": "https://www.hetzner.com/de/legal/legal-notice/",
  "scanStatus": "partial",
  "pagesAttempted": 20,
  "pagesSucceeded": 18,
  "pagesFailed": 2,
  "crawlStopReason": "max_pages",
  "crawledAt": "2026-09-26T18:28:18.304304+00:00",
  "emails": [
    "info@hetzner.com"
  ],
  "phones": [
    "+4998315050"
  ],
  "contactEvidence": [
    {
      "type": "email",
      "value": "info@hetzner.com",
      "platform": null,
      "emailDomainRelation": "same_site_domain",
      "sources": [
        {
          "sourceUrl": "https://www.hetzner.com/de/legal/legal-notice/",
          "methods": [
            "mailto",
            "text"
          ],
          "rawValues": [
            "info@hetzner.com"
          ],
          "fetchedAt": "2026-09-26T18:26:48.011348+00:00",
          "pageSha256": "9123f71fbc272b026f63032caa0a6d6b485f2fe2ebf37e70e2a3212dcb74ea5a"
        }
      ]
    }
  ]
}
```

`same_site_domain` is a domain-string relationship, not proof that the address is monitored or belongs to a particular employee. Use `emailDomainFilter: "same-site"` if you only want addresses matching the input site's host or its email subdomains. The default, `all`, also retains accepted third-party addresses explicitly published on the pages.

### All input options

| Field | Default | Meaning |
|---|---|---|
| `urls` | `[]` | Up to 10,000 nonempty strings: HTTP(S) URLs or bare domains. Bare domains use HTTPS. Repeated normalized inputs remain separate output tasks. |
| `startUrl` | omitted | Legacy single URL string, used only when `urls` is absent or empty. |
| `maxPagesPerSite` | `20` | Integer 1–200; attempted page URLs per crawl, including the starting page. Robots requests, redirects and retries add HTTP requests. |
| `maxConcurrency` | `5` | Integer 1–20; page concurrency per website. All websites share a ceiling of 20 active HTTP requests. |
| `maxWebsitesConcurrency` | `5` | Integer 1–20; concurrent input workers, subject to the shared HTTP ceiling and host delays. |
| `requestTimeoutSecs` | `15` | Integer 5–60; timeout per HTTP request, including its response body. |
| `requestDelayMillis` | `1000` | Integer 0–60000; minimum spacing between request starts to a host. Robots rules can increase it. |
| `maxRetries` | `1` | Integer 0–2; transient errors, HTTP 429 and selected 5xx responses. Does not bypass access restrictions. |
| `maxResults` | `10000` | Integer 1–10000; maximum **total saved rows, including failed rows**. Some inputs may already be in flight when a limit is reached. |
| `defaultPhoneRegion` | omitted | Supported uppercase country code, such as `DE`. Otherwise use a country TLD or explicit HTML language-region. Plain `en` does not imply US. |
| `emailDomainFilter` | `"all"` | `all` or `same-site`. Matching uses the input host without `www` and its email subdomains, not a company-ownership database. |
| `verifyEmails` | `false` | Enables the screening level below. When false, no email-verification MX or SMTP operations occur. Website DNS still occurs. |
| `verificationLevel` | `"mx"` | `format`, `mx` or `smtp`. Only active with `verifyEmails: true`. SMTP is experimental; it never certifies a mailbox. |
| `dnsResolver` | `"system"` | `system` or `google-doh`. The latter explicitly uses Google Public DNS over HTTPS for IPv4 website lookup and discloses the queried hostname to Google. Both reject nonpublic targets. |
| `useProxy` | `false` | Enables your own Apify proxy access or custom public HTTP(S) proxy. Does not purchase proxy access. |
| `proxyConfiguration` | omitted | SDK proxy settings: `useApifyProxy`, `apifyProxyGroups`, `apifyProxyCountry`, `apifyProxySubdivision`, or `proxyUrls`. Ignored when `useProxy` is false; invalid enabled configuration fails explicitly. |

Unknown input fields and invalid types are rejected before collection. Empty input returns no rows. Unsafe URLs rejected during normalization appear in diagnostics rather than the dataset, without echoing their raw values. A later DNS or fetch safety rejection can produce a safe failed record without contacting the blocked destination.

`maxTotalChargeUsd` is a run option, **not** an input field. Use the run's event-cost limit or the client's `max_total_charge_usd` argument. Do not put it inside the input JSON.

### Dataset fields and run results

The default dataset has one row per processed valid input, including repeated URLs and complete scan failures, unless a limit or interruption prevents processing or saving it.

| Dataset field | Meaning |
|---|---|
| `websiteUrl` | Normalized input URL. Final page URLs appear in `sourcePages`. |
| `emails`, `phones` | Sorted, deduplicated strings within the row. Phones use E.164, with extensions retained when present. Empty arrays mean no accepted values were found. |
| `socialLinks` | Selected link or null for each of `linkedin`, `twitter`, `facebook`, `instagram`, `youtube`, `github`, `tiktok`. Twitter/X links use the `twitter` key. |
| `socialProfiles` | All accepted profile URLs per platform, including multiple channels. No cross-site ownership enrichment is performed. |
| `socialRepositories` | Accepted published GitHub repository and repository-resource URLs, under `github`. |
| `socialLinkTypes` | `profile`, `repository` or null for each selected social link. GitHub prefers an observed profile; only when none exists does it select the first sorted repository. |
| `contactPageUrl`, `contactPageUrls` | Discovered contact-page URLs; discovery does not guarantee accessibility. The singular value prefers a fetched page. |
| `scanStatus` | `succeeded`: all attempted pages readable; `partial`: some readable and some failed; `failed`: none readable. Success does not mean the whole site was crawled. |
| `failureReason` | First safe failure category/message, or null. Detailed failures are in `SOURCE_DIAGNOSTICS`. |
| `pagesAttempted`, `pagesSucceeded`, `pagesFailed` | Logical page counts: attempted = succeeded + failed. |
| `pagesCrawled` | Compatibility alias for **attempted**, not successful, pages. |
| `crawlStopReason`, `queuedPagesRemaining` | `queue_exhausted`, `max_pages` or `budget_limit`, plus the remaining discovered queue size. |
| `crawledAt` | UTC record completion time. Cached duplicate copies retain the original timestamp. |
| `sourcePages` | Readable pages with requested/final URL, HTTP status, fetch time, SHA-256, byte count, title and phone-region hint. Full HTML is not included. |
| `hasContacts` | Whether any accepted email, phone or social link was found. |
| `contactEvidence` | Type, normalized value, platform when applicable, and all observed source URLs with methods, raw values, fetch times and page hashes. Email entries include `emailDomainRelation`; social entries include `linkType`. |
| `emailVerification` | Optional per-email screening results, present only when enabled on a readable record. Details below. |

The `OUTPUT` summary separates `pushedRows` (all saved rows), `readableRows` and `failedRows`. `duplicateInputs`, `cacheHits` and `uniqueCrawlsStarted` describe reuse; legacy `duplicatesSkipped` remains zero. `processedWebsites` includes diagnostic entries, while `failedWebsites` counts failed input diagnostics and may differ from saved failed rows.

Coverage and delivery fields include `requestedUrls`, `unstartedWebsites`, `unprocessedInputIndexes`, `notSavedInputIndexes`, `budgetStopped`, `limitReason`, `httpRequests`, `httpRetries`, `durationSeconds` and `finishedAt`. Indexes are zero-based. `requestedPageConcurrency`, `requestedWebsiteConcurrency`, `effectiveGlobalConcurrency`, `requestRouting` and `emptyInput` describe the run settings. `SOURCE_DIAGNOSTICS` carries each processed input's index, scan outcome, cache use, page failures, delivery outcome and termination details.

`OUTPUT.status` can be `succeeded`, `succeeded_empty`, `partial`, `failed` or `budget_limit_reached`. Input, proxy or pricing initialization errors use `invalid_input`, `proxy_configuration_failed` or `pricing_configuration_failed` with a message; normal summary fields may be absent. An all-failed scan saves its failure rows and diagnostics, then fails the Actor run. A mixed scan can have platform status `SUCCEEDED` with application status `partial`—read both.

### Repeated inputs, failures and limits

Repeated normalized URLs share an in-flight crawl and reuse a readable result within the same run. Each input gets an independent row copy. Concurrent duplicates can share a failure; a failed result is not kept for later duplicates, which may attempt it again. The cache does not persist across runs. Rows arrive in completion order; restart/resume does not provide exactly-once delivery across runs.

A readable site with no contacts still produces a successful or partial row. A wholly failed attempted site produces a row with `scanStatus: "failed"`, empty contacts and a safe reason. Inputs never attempted because of a budget, row limit or cancellation do not receive invented failure rows.

Budget and row limits stop new scheduling. Already-started work may appear as unsaved in diagnostics. An in-flight complete failure may still be saved without a website event after the successful-result event budget is exhausted, within `maxResults`. A cancellation can leave the run without a final summary. Inspect the dataset and delivery indexes before resubmitting unfinished inputs.

#### Real duplicate and 404 example

This exact input was run on **2026-09-27 UTC (2026-09-28 Asia/Shanghai)** with version **0.2.1**, build `oN3gtus37wTaV6Jb4`, run `HMXJzyvaMxsKgd2mz`:

```json
{
  "urls": [
    "https://example.com/",
    "https://example.com/",
    "https://example.com/website-contacts-missing-page-20260926"
  ],
  "maxPagesPerSite": 1,
  "maxConcurrency": 1,
  "maxWebsitesConcurrency": 3,
  "maxResults": 3,
  "requestTimeoutSecs": 15,
  "maxRetries": 0,
  "useProxy": false,
  "verifyEmails": false,
  "verificationLevel": "mx",
  "requestDelayMillis": 1000,
  "dnsResolver": "system",
  "emailDomainFilter": "all"
}
```

The saved dataset excerpt below preserves the order and values of all three rows, with other fields omitted. The first two complete saved records were identical, including their timestamps. This is a delivery/failure example with **zero contacts**; it does not demonstrate contact recall.

```json
[
  {
    "websiteUrl": "https://example.com/",
    "emails": [],
    "phones": [],
    "scanStatus": "succeeded",
    "failureReason": null,
    "pagesAttempted": 1,
    "pagesSucceeded": 1,
    "pagesFailed": 0,
    "crawledAt": "2026-09-27T18:34:34.582482+00:00"
  },
  {
    "websiteUrl": "https://example.com/",
    "emails": [],
    "phones": [],
    "scanStatus": "succeeded",
    "failureReason": null,
    "pagesAttempted": 1,
    "pagesSucceeded": 1,
    "pagesFailed": 0,
    "crawledAt": "2026-09-27T18:34:34.582482+00:00"
  },
  {
    "websiteUrl": "https://example.com/website-contacts-missing-page-20260926",
    "emails": [],
    "phones": [],
    "scanStatus": "failed",
    "failureReason": {
      "category": "HTTP_ERROR",
      "message": "The website returned HTTP 404"
    },
    "pagesAttempted": 1,
    "pagesSucceeded": 0,
    "pagesFailed": 1,
    "crawledAt": "2026-09-27T18:34:35.572753+00:00"
  }
]
```

The saved summary reported `status: "partial"`, `pushedRows: 3`, `readableRows: 2`, `failedRows: 1`, `cacheHits: 1`, `uniqueCrawlsStarted: 2` and `httpRequests: 3` (including robots). All inputs were saved. `budgetStopped: true` and `limitReason: "max_results"` record that the three-row cap was reached; they do not imply missing rows when both unfinished-index arrays are empty. This run was unpriced: `billingEnabledByThisPackage: false`, `billingMode: "free"`, `websiteEventPriceUsd: null`. It does not establish a paid price or a paid cloud result.

### Costs and billing

Check the Actor's current Pricing tab before running. When website-event pricing is active, the billing unit is **one saved readable website row** (`website-scanned`), not an email, phone, page or HTTP request. A partial-readable row, a readable row with no contacts and each saved readable duplicate count as one website event. A complete failure row has **zero website-event fee**. That statement does not promise zero platform usage or proxy cost; check the platform terms and your proxy plan.

Supported modes are unpriced/FREE, or website-event pricing with a zero effective automatic dataset-item price. A positive automatic dataset-item price would charge failed rows, so that configuration is rejected before collection. Pay-per-event configuration without `website-scanned` and other unsupported pricing models also fail explicitly. The Actor does not change its prices during a run.

`OUTPUT.billingEnabledByThisPackage` means this run has an effective website-event configuration, not that a positive dollar charge occurred. `billingMode` is `free` or `website_event` in Actor runs (`local_unbilled` is reserved for standalone execution). `configuredChargeEvent` is `website-scanned` or null; `websiteEventPriceUsd` is its decimal-string unit price or null. A zero-price website event reports enabled/`website_event` with a zero price. Use the platform's actual charge records for money totals; event counts alone are not dollars.

An event-cost limit governs paid result delivery; `maxResults` governs all saved rows. A paid cap too small for one readable row can stop the run before any input is attempted, including inputs that might have failed. An unpriced run does not treat a nominal zero event budget as a reason to stop.

#### Current event prices

| Account discount tier | Per 1,000 saved readable website records |
|---|---:|
| FREE | $0.80 |
| BRONZE | $0.70 |
| SILVER | $0.60 |
| GOLD | $0.40 |
| PLATINUM | $0.40 |
| DIAMOND | $0.40 |

The automatic platform start event is **$0.00005 per start for up to 1 GB RAM**; larger allocations trigger more start events. **Apify platform compute, storage and network usage are charged additionally.** There is no positive automatic dataset-item fee. These are result-event prices, not an all-inclusive job quote. Optional user-provided proxies can add their own fees. Check the Pricing tab for your effective account tier.

On 2026-09-30 UTC, a private billing check on build `0.2.1` saved two readable duplicate example.com rows and one real HTTP-404 failure row. Platform run `sIggnG6gB2kjIOdDU` succeeded; the application summary correctly remained `partial`. The platform reported two `website-scanned` events at $0.0007 each and one automatic start event at $0.00005: **$0.00145 in nominal event fees, plus platform usage**. The failure row incurred no website event. This zero-contact sample verifies delivery and charging; it does not measure contact recall or large-job costs.

### Optional email screening

Leave `verifyEmails` false for extraction only. To add MX checks, use:

```json
{
  "urls": ["https://www.hetzner.com/de/legal/legal-notice/"],
  "maxPagesPerSite": 1,
  "verifyEmails": true,
  "verificationLevel": "mx"
}
```

| Level | What it tells you | What it does not establish |
|---|---|---|
| `format` | Accepted email syntax. | Mail routing, mailbox existence or delivery. |
| `mx` | Syntax and published DNS MX records, cached by domain within the run. | Whether the mailbox exists or accepts mail. |
| `smtp` | Format → MX → an optional, bounded port-25 envelope observation for eligible routes. | Verified mailbox existence, catch-all behavior or deliverability. Live SMTP has not been validated. |

Verification output includes `email`, `isValidFormat`, `hasMxRecords`, `isVerified`, `confidenceScore`, `isDisposable`, `isFreeProvider`, `isRoleAccount`, `provider`, `verificationLevel`, `verificationStatus`, `disposableCheckCoverage` and `mailboxExistence`. `isVerified` always remains false. Scores are simple heuristics (30 for syntax, 45 with MX, zero for invalid/disposable), not probabilities. Provider/role/disposable lists have limited coverage.

`hasMxRecords` is true, false or null (not checked, unavailable or invalid data). A sole null MX (`0 .`) declares no mail service; mixed or invalid null-MX records remain uncertain and are not probed. Absent MX records do not trigger an implicit A/AAAA mail-delivery fallback. Neither missing MX nor a low score should be treated as a complete deliverability verdict.

**SMTP is experimental and optional.** It sends EHLO, an empty `MAIL FROM:<>` and one RCPT for an extracted address, then attempts bounded RSET/QUIT cleanup. It never sends DATA, a message body, AUTH, credentials, VRFY or guessed catch-all recipients. It uses plain TCP; STARTTLS and SMTPUTF8 are not implemented. Servers requiring encryption or authentication yield an unknown observation.

The additional `smtp` object reports `status: "unknown"`, `reason`, `stage`, `replyCode`, `attempted` and cleanup results. `attempted` can be null when progress is unknown after timeout. RCPT acceptance is `rcpt_accepted_unconfirmed`, not proof of existence or delivery. Forwarding/cannot-verify replies, 4xx/5xx, timeouts, blocked ports and malformed dialogues remain diagnostic observations. `catchAllStatus: "not_checked"`, `deliverability: "not_tested"` and `mailboxExistence: "unknown"` remain explicit.

Known Google, Microsoft and Yahoo recipient/MX routes are skipped as unreliable for probes. Other providers can also block or mislead them. SMTP destinations must resolve entirely to public addresses, and the connection uses the checked IP. SMTP uses direct TCP and system DNS, separately from the website HTTP proxy and DNS-over-HTTPS options. Full server dialogues and session identifiers are not retained.

SMTP has independent caps: two active probes, two-second start spacing per recipient domain/shared MX, ten reservations per domain and 100 per run, at most two MX connection attempts per reservation. Probe deadlines include waiting and cleanup: 12 seconds total, three seconds per operation, 18 seconds including verification/MX. MX has four slots and a five-second deadline. Reply limits are 16 lines, 512 bytes per line and 16,384 bytes per session. This protocol and its safety limits have deterministic offline coverage; live SMTP acceptance remains outstanding.

### Collection boundaries and proxy use

Contact-related links take priority, followed by about/company/support pages and remaining internal HTML links. The crawler recognizes several language-specific contact labels, follows actual query-page links, removes fragments/tracking parameters and queues each discovered page once. It does not invent contact paths or discover pages through sitemaps.

Scope is the exact input host and its `www` alias. Other subdomains and cross-domain redirects are excluded. Robots rules are checked per origin and redirect destination; denied or unavailable robots policy fails closed. Private/local addresses, credential-bearing URLs, sensitive token parameters and nonstandard target ports are rejected.

Email sources include mailto links, visible text, explicit at/dot obfuscation and Organization JSON-LD. Person JSON-LD, script/style/template content and hidden attributes are excluded. Common placeholders, noreply addresses and asset-like matches are filtered. Phone parsing rejects known fictional/example patterns and selected reserved ranges; a plausible number is not proof of assignment or reachability. Social filtering accepts company/school LinkedIn pages and supported profile/repository links, but coverage and representative-link choices can differ across websites.

There is no JavaScript/browser rendering, login, browser-cookie access, form submission, CAPTCHA bypass, OCR or PDF extraction. Contacts available only through those mechanisms may be absent. A discovered contact URL may be blocked or outside the page budget. GitHub links appear only if they were actually present on fetched pages; the Actor does not guess repositories.

For proxy access you already have, a minimal opt-in input is:

```json
{
  "urls": ["https://www.scrapingbee.com/"],
  "maxPagesPerSite": 3,
  "useProxy": true,
  "proxyConfiguration": {"useApifyProxy": true}
}
```

Initialization and routing failures are explicit; there is no silent direct fallback. Target and proxy addresses are checked, and the proxy is asked to connect to the validated numeric target IP while preserving the website Host/TLS name. Some providers reject numeric-IP CONNECT routes. Live paid Apify-proxy routing has not been validated. Proxy providers are trusted transports and their charges depend on your access plan. HTTP proxy settings do not route SMTP.

### Migration from Website Contact Finder

This Actor accepts the common `automation-lab/website-contact-finder` inputs: `urls`, `maxPagesPerSite`, `maxConcurrency`, `maxWebsitesConcurrency`, `requestTimeoutSecs`, `useProxy`, `proxyConfiguration`, `verifyEmails` and `verificationLevel`. It also supports legacy `startUrl` and the controls listed above. Duplicate inputs produce separate rows, and wholly failed scans are retained without a website-event fee.

Retained core fields include `websiteUrl`, `emails`, `phones`, `socialLinks`, `contactPageUrl`, `pagesCrawled`, `scanStatus`, `failureReason`, the attempted/succeeded/failed page counts and `crawledAt`. Additional evidence, all-profile arrays and typed repository links support consumers that need more context.

This is **not a full drop-in replacement**. Keep these differences in your integration:

| Area | Migration consideration |
|---|---|
| Coverage | Static HTML, host scope, robots rules, available links, page limits and filters affect which contacts are returned. No all-sites or permanent success guarantee. |
| Phones/socials | Phones are normalized to E.164. A representative social link may differ; use all-profile/repository arrays and source evidence. Links do not prove ownership. |
| Email screening | Format/MX heuristics and provider flags can differ. SMTP observations are not mailbox or catch-all verification; live SMTP remains unvalidated. |
| Delivery | Completion order is not input order. Failed rows stay in the dataset; all-failed runs fail after saving diagnostics. The cache is run-local. |
| Concurrency/proxy | Requested concurrency up to 20 shares a 20-request ceiling. Live high-volume throughput and paid proxy routing have not been established. |

### Run from Python and automate exports

Install `apify-client` in your client environment. Set `APIFY_TOKEN` securely in that environment; never put it in the input or source URLs. The example uses Actor ID `OOjMOcPufmpGn0fhg`. If desired, set `APIFY_MAX_TOTAL_CHARGE_USD` to your chosen event-cost cap before execution; it is a budget, not a quoted product price.

This client example uses the Actor's default build. For version-specific behavior, pass an available build number using the client's `build` argument; a documented version does not by itself change the Actor's default.

```python
import json
import os
from decimal import Decimal

from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
options = {}
if os.environ.get("APIFY_MAX_TOTAL_CHARGE_USD"):
    options["max_total_charge_usd"] = Decimal(
        os.environ["APIFY_MAX_TOTAL_CHARGE_USD"]
    )

run = client.actor("OOjMOcPufmpGn0fhg").call(
    run_input={
        "urls": ["https://www.scrapingbee.com/"],
        "maxPagesPerSite": 3,
        "maxConcurrency": 1,
        "maxWebsitesConcurrency": 1,
        "verifyEmails": False,
        "useProxy": False,
    },
    timeout_secs=300,
    **options,
)
if run is None:
    raise RuntimeError("No run record returned")

## Read persisted rows even when the run failed after an all-failed scan.
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(json.dumps(row, ensure_ascii=False))

store = client.key_value_store(run["defaultKeyValueStoreId"])
for key in ("OUTPUT", "SOURCE_DIAGNOSTICS"):
    record = store.get_record(key)
    print(key, json.dumps(record["value"] if record else None, ensure_ascii=False))
print("Run status:", run["status"])
```

For recurring refreshes, save the input as an Apify task and schedule it. Keep each run ID with your exports, filter or route failed rows separately, and deduplicate downstream across runs if needed. A fresh run performs a fresh acquisition; prior runs do not populate its cache.

### FAQ

**Why did a successful scan return no contacts?** A readable page may publish none, link to contacts beyond your page limit or require JavaScript. Check the attempted/readable counts, `sourcePages`, contact-page discovery and stop reason before increasing coverage.

**Are all extracted addresses contacts for the target company?** No. A site can publish hosting-provider, legal or other third-party addresses. Review `emailDomainRelation` and source context, or select `same-site`. Domain matching still does not prove ownership.

**Why are repeated URLs charged again in paid mode?** The unit is a saved readable input row. The run-local cache reduces repeated acquisition, while each delivered readable copy is a separate website event. Remove duplicates from your input if you only need one row.

**Can I use SMTP results as a validated mailing list?** No. Neither syntax, MX records nor RCPT acceptance proves delivery. SMTP remains experimental and never returns a verified-mailbox claim.

**Why can a failed run still have results?** Complete failures are useful records. All-failed runs save them before failing; mixed runs can succeed on the platform while `OUTPUT.status` is `partial`. Read the dataset and diagnostics instead of relying only on the run status.

# Actor input Schema

## `urls` (type: `array`):

HTTP(S) URLs or domains. Each processed valid input retains its own result, including repeated inputs and failed scans. Limits may leave inputs unprocessed. Empty input returns no rows.

## `startUrl` (type: `string`):

Used only if urls is missing or empty.

## `maxPagesPerSite` (type: `integer`):

Includes the starting page; contact and about pages are prioritized. Robots and redirects are extra HTTP requests.

## `maxConcurrency` (type: `integer`):

Honor up to 20 per site, subject to a run-wide ceiling of 20 active HTTP requests and per-host spacing.

## `maxWebsitesConcurrency` (type: `integer`):

Honor 1–20 websites concurrently; their HTTP requests share a global ceiling of 20.

## `requestTimeoutSecs` (type: `integer`):

Finite timeout per HTTP request, including response body.

## `useProxy` (type: `boolean`):

Opt in to your own Apify proxy access or custom public HTTP(S) proxies. No proxy is purchased. Initialization failures are explicit.

## `proxyConfiguration` (type: `object`):

Passed to the Apify SDK when useProxy=true. Supports Apify groups, country/subdivision and custom proxyUrls. With useProxy=false no proxy is used.

## `verifyEmails` (type: `boolean`):

Opt in to format/MX checks and optional bounded SMTP envelope observations. Mailbox existence and deliverability are never guaranteed.

## `verificationLevel` (type: `string`):

format: syntax; mx: syntax+MX; smtp: additionally bounded port-25 envelope observations, never DATA/AUTH/mail. Known unreliable Google/Microsoft/Yahoo routes are skipped. RCPT acceptance remains unknown for existence; catch-all is not checked. SMTP has offline protocol tests only; live SMTP is unvalidated.

## `defaultPhoneRegion` (type: `string`):

Optional uppercase ISO country, e.g. DE or US. Otherwise use site country TLD or explicit HTML language region; no default US guessing.

## `requestDelayMillis` (type: `integer`):

Robots crawl-delay and request-rate can increase this delay.

## `maxRetries` (type: `integer`):

Retry timeouts, transient network errors, 429 and selected 5xx; never bypass access blocks.

## `maxResults` (type: `integer`):

Stops scheduling when reached. Up to maxWebsitesConcurrency sites can already be in flight.

## `dnsResolver` (type: `string`):

system fails closed on local/private DNS answers. Explicit google-doh uses pinned public resolver 8.8.8.8 over verified HTTPS, resolves IPv4, and still rejects nonpublic target addresses.

## `emailDomainFilter` (type: `string`):

all preserves every accepted publicly published email, including service-provider contacts. same-site limits emails to the site host (excluding www) and its email subdomains. This is a string domain relation, not an ownership or employment claim.

## Actor input object example

```json
{
  "urls": [
    "https://www.scrapingbee.com/"
  ],
  "maxPagesPerSite": 3,
  "maxConcurrency": 5,
  "maxWebsitesConcurrency": 5,
  "requestTimeoutSecs": 15,
  "useProxy": false,
  "verifyEmails": false,
  "verificationLevel": "mx",
  "requestDelayMillis": 1000,
  "maxRetries": 1,
  "maxResults": 10000,
  "dnsResolver": "system",
  "emailDomainFilter": "all"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `diagnostics` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.scrapingbee.com/"
    ],
    "maxPagesPerSite": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("peerless_columbine/website-contact-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.scrapingbee.com/"],
    "maxPagesPerSite": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("peerless_columbine/website-contact-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.scrapingbee.com/"
  ],
  "maxPagesPerSite": 3
}' |
apify call peerless_columbine/website-contact-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,peerless_columbine/website-contact-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OOjMOcPufmpGn0fhg/builds/nxeMwyqsUYVbJaF8O/openapi.json
