Website Contact Finder | Emails, Phones & Source Evidence
Pricing
from $0.40 / 1,000 readable website records
Website Contact Finder | Emails, Phones & Source Evidence
Extract public business emails, phones and social profiles from websites. Follow contact pages, retain source evidence, preserve duplicate inputs and failure records. Clear page limits and optional email checks.
Pricing
from $0.40 / 1,000 readable website records
Rating
0.0
(0)
Developer
tingyou333 zhuang
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Website Contacts | Emails, Phones & Source URLs
Turn a list of company websites into records of their publicly listed email addresses, phone numbers and social links. Every accepted contact includes its source page, extraction method, fetch time and page hash, so you can check the context before importing it into a CRM.
Use this Actor to review company contact pages, refresh published support channels or audit a website list. It reads public HTML and follows contact, about and company links on the same site. It does not guess email addresses, search for people or infer relationships across websites. A published address or social link can belong to a third party; its presence does not establish company ownership.
Start with a small scan
Paste this input into the Actor's JSON input editor and run it:
{"urls": ["https://www.scrapingbee.com/"],"maxPagesPerSite": 3,"maxConcurrency": 1,"maxWebsitesConcurrency": 1,"verifyEmails": false,"useProxy": false}
Open the default dataset to view or export website records. Keep scanStatus when importing results: a readable page with no contacts and a failed scan both have empty contact arrays, but mean different things. The run's key-value store contains OUTPUT for totals and limits, and SOURCE_DIAGNOSTICS for individual input failures and delivery details.
Three pages are a useful small starting point, not a promise of complete contact coverage. Increase maxPagesPerSite for deeper discovery or start directly at a public contact/legal page. Omitting this option uses 20 pages. Source content and access can change between runs.
What a contact record looks like
This is an excerpt from an actual Hetzner scan on 2026-09-26 UTC, version 0.1.3, build mGDQUJDwzRxITNXbg, run tqMJFBi3BnroHGP8s. It preserves the recorded values and shows only one email, one phone and that email's complete evidence entry; other contacts and fields are omitted. Eighteen pages were readable and two failed, so the row is correctly marked partial. This older sample illustrates extraction and provenance; it is not a fresh scan or a claim about every website.
{"websiteUrl": "https://www.hetzner.com/de/","contactPageUrl": "https://www.hetzner.com/de/legal/legal-notice/","scanStatus": "partial","pagesAttempted": 20,"pagesSucceeded": 18,"pagesFailed": 2,"crawlStopReason": "max_pages","crawledAt": "2026-09-26T18:28:18.304304+00:00","emails": ["info@hetzner.com"],"phones": ["+4998315050"],"contactEvidence": [{"type": "email","value": "info@hetzner.com","platform": null,"emailDomainRelation": "same_site_domain","sources": [{"sourceUrl": "https://www.hetzner.com/de/legal/legal-notice/","methods": ["mailto","text"],"rawValues": ["info@hetzner.com"],"fetchedAt": "2026-09-26T18:26:48.011348+00:00","pageSha256": "9123f71fbc272b026f63032caa0a6d6b485f2fe2ebf37e70e2a3212dcb74ea5a"}]}]}
same_site_domain is a domain-string relationship, not proof that the address is monitored or belongs to a particular employee. Use emailDomainFilter: "same-site" if you only want addresses matching the input site's host or its email subdomains. The default, all, also retains accepted third-party addresses explicitly published on the pages.
All input options
| Field | Default | Meaning |
|---|---|---|
urls | [] | Up to 10,000 nonempty strings: HTTP(S) URLs or bare domains. Bare domains use HTTPS. Repeated normalized inputs remain separate output tasks. |
startUrl | omitted | Legacy single URL string, used only when urls is absent or empty. |
maxPagesPerSite | 20 | Integer 1–200; attempted page URLs per crawl, including the starting page. Robots requests, redirects and retries add HTTP requests. |
maxConcurrency | 5 | Integer 1–20; page concurrency per website. All websites share a ceiling of 20 active HTTP requests. |
maxWebsitesConcurrency | 5 | Integer 1–20; concurrent input workers, subject to the shared HTTP ceiling and host delays. |
requestTimeoutSecs | 15 | Integer 5–60; timeout per HTTP request, including its response body. |
requestDelayMillis | 1000 | Integer 0–60000; minimum spacing between request starts to a host. Robots rules can increase it. |
maxRetries | 1 | Integer 0–2; transient errors, HTTP 429 and selected 5xx responses. Does not bypass access restrictions. |
maxResults | 10000 | Integer 1–10000; maximum total saved rows, including failed rows. Some inputs may already be in flight when a limit is reached. |
defaultPhoneRegion | omitted | Supported uppercase country code, such as DE. Otherwise use a country TLD or explicit HTML language-region. Plain en does not imply US. |
emailDomainFilter | "all" | all or same-site. Matching uses the input host without www and its email subdomains, not a company-ownership database. |
verifyEmails | false | Enables the screening level below. When false, no email-verification MX or SMTP operations occur. Website DNS still occurs. |
verificationLevel | "mx" | format, mx or smtp. Only active with verifyEmails: true. SMTP is experimental; it never certifies a mailbox. |
dnsResolver | "system" | system or google-doh. The latter explicitly uses Google Public DNS over HTTPS for IPv4 website lookup and discloses the queried hostname to Google. Both reject nonpublic targets. |
useProxy | false | Enables your own Apify proxy access or custom public HTTP(S) proxy. Does not purchase proxy access. |
proxyConfiguration | omitted | SDK proxy settings: useApifyProxy, apifyProxyGroups, apifyProxyCountry, apifyProxySubdivision, or proxyUrls. Ignored when useProxy is false; invalid enabled configuration fails explicitly. |
Unknown input fields and invalid types are rejected before collection. Empty input returns no rows. Unsafe URLs rejected during normalization appear in diagnostics rather than the dataset, without echoing their raw values. A later DNS or fetch safety rejection can produce a safe failed record without contacting the blocked destination.
maxTotalChargeUsd is a run option, not an input field. Use the run's event-cost limit or the client's max_total_charge_usd argument. Do not put it inside the input JSON.
Dataset fields and run results
The default dataset has one row per processed valid input, including repeated URLs and complete scan failures, unless a limit or interruption prevents processing or saving it.
| Dataset field | Meaning |
|---|---|
websiteUrl | Normalized input URL. Final page URLs appear in sourcePages. |
emails, phones | Sorted, deduplicated strings within the row. Phones use E.164, with extensions retained when present. Empty arrays mean no accepted values were found. |
socialLinks | Selected link or null for each of linkedin, twitter, facebook, instagram, youtube, github, tiktok. Twitter/X links use the twitter key. |
socialProfiles | All accepted profile URLs per platform, including multiple channels. No cross-site ownership enrichment is performed. |
socialRepositories | Accepted published GitHub repository and repository-resource URLs, under github. |
socialLinkTypes | profile, repository or null for each selected social link. GitHub prefers an observed profile; only when none exists does it select the first sorted repository. |
contactPageUrl, contactPageUrls | Discovered contact-page URLs; discovery does not guarantee accessibility. The singular value prefers a fetched page. |
scanStatus | succeeded: all attempted pages readable; partial: some readable and some failed; failed: none readable. Success does not mean the whole site was crawled. |
failureReason | First safe failure category/message, or null. Detailed failures are in SOURCE_DIAGNOSTICS. |
pagesAttempted, pagesSucceeded, pagesFailed | Logical page counts: attempted = succeeded + failed. |
pagesCrawled | Compatibility alias for attempted, not successful, pages. |
crawlStopReason, queuedPagesRemaining | queue_exhausted, max_pages or budget_limit, plus the remaining discovered queue size. |
crawledAt | UTC record completion time. Cached duplicate copies retain the original timestamp. |
sourcePages | Readable pages with requested/final URL, HTTP status, fetch time, SHA-256, byte count, title and phone-region hint. Full HTML is not included. |
hasContacts | Whether any accepted email, phone or social link was found. |
contactEvidence | Type, normalized value, platform when applicable, and all observed source URLs with methods, raw values, fetch times and page hashes. Email entries include emailDomainRelation; social entries include linkType. |
emailVerification | Optional per-email screening results, present only when enabled on a readable record. Details below. |
The OUTPUT summary separates pushedRows (all saved rows), readableRows and failedRows. duplicateInputs, cacheHits and uniqueCrawlsStarted describe reuse; legacy duplicatesSkipped remains zero. processedWebsites includes diagnostic entries, while failedWebsites counts failed input diagnostics and may differ from saved failed rows.
Coverage and delivery fields include requestedUrls, unstartedWebsites, unprocessedInputIndexes, notSavedInputIndexes, budgetStopped, limitReason, httpRequests, httpRetries, durationSeconds and finishedAt. Indexes are zero-based. requestedPageConcurrency, requestedWebsiteConcurrency, effectiveGlobalConcurrency, requestRouting and emptyInput describe the run settings. SOURCE_DIAGNOSTICS carries each processed input's index, scan outcome, cache use, page failures, delivery outcome and termination details.
OUTPUT.status can be succeeded, succeeded_empty, partial, failed or budget_limit_reached. Input, proxy or pricing initialization errors use invalid_input, proxy_configuration_failed or pricing_configuration_failed with a message; normal summary fields may be absent. An all-failed scan saves its failure rows and diagnostics, then fails the Actor run. A mixed scan can have platform status SUCCEEDED with application status partial—read both.
Repeated inputs, failures and limits
Repeated normalized URLs share an in-flight crawl and reuse a readable result within the same run. Each input gets an independent row copy. Concurrent duplicates can share a failure; a failed result is not kept for later duplicates, which may attempt it again. The cache does not persist across runs. Rows arrive in completion order; restart/resume does not provide exactly-once delivery across runs.
A readable site with no contacts still produces a successful or partial row. A wholly failed attempted site produces a row with scanStatus: "failed", empty contacts and a safe reason. Inputs never attempted because of a budget, row limit or cancellation do not receive invented failure rows.
Budget and row limits stop new scheduling. Already-started work may appear as unsaved in diagnostics. An in-flight complete failure may still be saved without a website event after the successful-result event budget is exhausted, within maxResults. A cancellation can leave the run without a final summary. Inspect the dataset and delivery indexes before resubmitting unfinished inputs.
Real duplicate and 404 example
This exact input was run on 2026-09-27 UTC (2026-09-28 Asia/Shanghai) with version 0.2.1, build oN3gtus37wTaV6Jb4, run HMXJzyvaMxsKgd2mz:
{"urls": ["https://example.com/","https://example.com/","https://example.com/website-contacts-missing-page-20260926"],"maxPagesPerSite": 1,"maxConcurrency": 1,"maxWebsitesConcurrency": 3,"maxResults": 3,"requestTimeoutSecs": 15,"maxRetries": 0,"useProxy": false,"verifyEmails": false,"verificationLevel": "mx","requestDelayMillis": 1000,"dnsResolver": "system","emailDomainFilter": "all"}
The saved dataset excerpt below preserves the order and values of all three rows, with other fields omitted. The first two complete saved records were identical, including their timestamps. This is a delivery/failure example with zero contacts; it does not demonstrate contact recall.
[{"websiteUrl": "https://example.com/","emails": [],"phones": [],"scanStatus": "succeeded","failureReason": null,"pagesAttempted": 1,"pagesSucceeded": 1,"pagesFailed": 0,"crawledAt": "2026-09-27T18:34:34.582482+00:00"},{"websiteUrl": "https://example.com/","emails": [],"phones": [],"scanStatus": "succeeded","failureReason": null,"pagesAttempted": 1,"pagesSucceeded": 1,"pagesFailed": 0,"crawledAt": "2026-09-27T18:34:34.582482+00:00"},{"websiteUrl": "https://example.com/website-contacts-missing-page-20260926","emails": [],"phones": [],"scanStatus": "failed","failureReason": {"category": "HTTP_ERROR","message": "The website returned HTTP 404"},"pagesAttempted": 1,"pagesSucceeded": 0,"pagesFailed": 1,"crawledAt": "2026-09-27T18:34:35.572753+00:00"}]
The saved summary reported status: "partial", pushedRows: 3, readableRows: 2, failedRows: 1, cacheHits: 1, uniqueCrawlsStarted: 2 and httpRequests: 3 (including robots). All inputs were saved. budgetStopped: true and limitReason: "max_results" record that the three-row cap was reached; they do not imply missing rows when both unfinished-index arrays are empty. This run was unpriced: billingEnabledByThisPackage: false, billingMode: "free", websiteEventPriceUsd: null. It does not establish a paid price or a paid cloud result.
Costs and billing
Check the Actor's current Pricing tab before running. When website-event pricing is active, the billing unit is one saved readable website row (website-scanned), not an email, phone, page or HTTP request. A partial-readable row, a readable row with no contacts and each saved readable duplicate count as one website event. A complete failure row has zero website-event fee. That statement does not promise zero platform usage or proxy cost; check the platform terms and your proxy plan.
Supported modes are unpriced/FREE, or website-event pricing with a zero effective automatic dataset-item price. A positive automatic dataset-item price would charge failed rows, so that configuration is rejected before collection. Pay-per-event configuration without website-scanned and other unsupported pricing models also fail explicitly. The Actor does not change its prices during a run.
OUTPUT.billingEnabledByThisPackage means this run has an effective website-event configuration, not that a positive dollar charge occurred. billingMode is free or website_event in Actor runs (local_unbilled is reserved for standalone execution). configuredChargeEvent is website-scanned or null; websiteEventPriceUsd is its decimal-string unit price or null. A zero-price website event reports enabled/website_event with a zero price. Use the platform's actual charge records for money totals; event counts alone are not dollars.
An event-cost limit governs paid result delivery; maxResults governs all saved rows. A paid cap too small for one readable row can stop the run before any input is attempted, including inputs that might have failed. An unpriced run does not treat a nominal zero event budget as a reason to stop.
Current event prices
| Account discount tier | Per 1,000 saved readable website records |
|---|---|
| FREE | $0.80 |
| BRONZE | $0.70 |
| SILVER | $0.60 |
| GOLD | $0.40 |
| PLATINUM | $0.40 |
| DIAMOND | $0.40 |
The automatic platform start event is $0.00005 per start for up to 1 GB RAM; larger allocations trigger more start events. Apify platform compute, storage and network usage are charged additionally. There is no positive automatic dataset-item fee. These are result-event prices, not an all-inclusive job quote. Optional user-provided proxies can add their own fees. Check the Pricing tab for your effective account tier.
On 2026-09-30 UTC, a private billing check on build 0.2.1 saved two readable duplicate example.com rows and one real HTTP-404 failure row. Platform run sIggnG6gB2kjIOdDU succeeded; the application summary correctly remained partial. The platform reported two website-scanned events at $0.0007 each and one automatic start event at $0.00005: $0.00145 in nominal event fees, plus platform usage. The failure row incurred no website event. This zero-contact sample verifies delivery and charging; it does not measure contact recall or large-job costs.
Optional email screening
Leave verifyEmails false for extraction only. To add MX checks, use:
{"urls": ["https://www.hetzner.com/de/legal/legal-notice/"],"maxPagesPerSite": 1,"verifyEmails": true,"verificationLevel": "mx"}
| Level | What it tells you | What it does not establish |
|---|---|---|
format | Accepted email syntax. | Mail routing, mailbox existence or delivery. |
mx | Syntax and published DNS MX records, cached by domain within the run. | Whether the mailbox exists or accepts mail. |
smtp | Format → MX → an optional, bounded port-25 envelope observation for eligible routes. | Verified mailbox existence, catch-all behavior or deliverability. Live SMTP has not been validated. |
Verification output includes email, isValidFormat, hasMxRecords, isVerified, confidenceScore, isDisposable, isFreeProvider, isRoleAccount, provider, verificationLevel, verificationStatus, disposableCheckCoverage and mailboxExistence. isVerified always remains false. Scores are simple heuristics (30 for syntax, 45 with MX, zero for invalid/disposable), not probabilities. Provider/role/disposable lists have limited coverage.
hasMxRecords is true, false or null (not checked, unavailable or invalid data). A sole null MX (0 .) declares no mail service; mixed or invalid null-MX records remain uncertain and are not probed. Absent MX records do not trigger an implicit A/AAAA mail-delivery fallback. Neither missing MX nor a low score should be treated as a complete deliverability verdict.
SMTP is experimental and optional. It sends EHLO, an empty MAIL FROM:<> and one RCPT for an extracted address, then attempts bounded RSET/QUIT cleanup. It never sends DATA, a message body, AUTH, credentials, VRFY or guessed catch-all recipients. It uses plain TCP; STARTTLS and SMTPUTF8 are not implemented. Servers requiring encryption or authentication yield an unknown observation.
The additional smtp object reports status: "unknown", reason, stage, replyCode, attempted and cleanup results. attempted can be null when progress is unknown after timeout. RCPT acceptance is rcpt_accepted_unconfirmed, not proof of existence or delivery. Forwarding/cannot-verify replies, 4xx/5xx, timeouts, blocked ports and malformed dialogues remain diagnostic observations. catchAllStatus: "not_checked", deliverability: "not_tested" and mailboxExistence: "unknown" remain explicit.
Known Google, Microsoft and Yahoo recipient/MX routes are skipped as unreliable for probes. Other providers can also block or mislead them. SMTP destinations must resolve entirely to public addresses, and the connection uses the checked IP. SMTP uses direct TCP and system DNS, separately from the website HTTP proxy and DNS-over-HTTPS options. Full server dialogues and session identifiers are not retained.
SMTP has independent caps: two active probes, two-second start spacing per recipient domain/shared MX, ten reservations per domain and 100 per run, at most two MX connection attempts per reservation. Probe deadlines include waiting and cleanup: 12 seconds total, three seconds per operation, 18 seconds including verification/MX. MX has four slots and a five-second deadline. Reply limits are 16 lines, 512 bytes per line and 16,384 bytes per session. This protocol and its safety limits have deterministic offline coverage; live SMTP acceptance remains outstanding.
Collection boundaries and proxy use
Contact-related links take priority, followed by about/company/support pages and remaining internal HTML links. The crawler recognizes several language-specific contact labels, follows actual query-page links, removes fragments/tracking parameters and queues each discovered page once. It does not invent contact paths or discover pages through sitemaps.
Scope is the exact input host and its www alias. Other subdomains and cross-domain redirects are excluded. Robots rules are checked per origin and redirect destination; denied or unavailable robots policy fails closed. Private/local addresses, credential-bearing URLs, sensitive token parameters and nonstandard target ports are rejected.
Email sources include mailto links, visible text, explicit at/dot obfuscation and Organization JSON-LD. Person JSON-LD, script/style/template content and hidden attributes are excluded. Common placeholders, noreply addresses and asset-like matches are filtered. Phone parsing rejects known fictional/example patterns and selected reserved ranges; a plausible number is not proof of assignment or reachability. Social filtering accepts company/school LinkedIn pages and supported profile/repository links, but coverage and representative-link choices can differ across websites.
There is no JavaScript/browser rendering, login, browser-cookie access, form submission, CAPTCHA bypass, OCR or PDF extraction. Contacts available only through those mechanisms may be absent. A discovered contact URL may be blocked or outside the page budget. GitHub links appear only if they were actually present on fetched pages; the Actor does not guess repositories.
For proxy access you already have, a minimal opt-in input is:
{"urls": ["https://www.scrapingbee.com/"],"maxPagesPerSite": 3,"useProxy": true,"proxyConfiguration": {"useApifyProxy": true}}
Initialization and routing failures are explicit; there is no silent direct fallback. Target and proxy addresses are checked, and the proxy is asked to connect to the validated numeric target IP while preserving the website Host/TLS name. Some providers reject numeric-IP CONNECT routes. Live paid Apify-proxy routing has not been validated. Proxy providers are trusted transports and their charges depend on your access plan. HTTP proxy settings do not route SMTP.
Migration from Website Contact Finder
This Actor accepts the common automation-lab/website-contact-finder inputs: urls, maxPagesPerSite, maxConcurrency, maxWebsitesConcurrency, requestTimeoutSecs, useProxy, proxyConfiguration, verifyEmails and verificationLevel. It also supports legacy startUrl and the controls listed above. Duplicate inputs produce separate rows, and wholly failed scans are retained without a website-event fee.
Retained core fields include websiteUrl, emails, phones, socialLinks, contactPageUrl, pagesCrawled, scanStatus, failureReason, the attempted/succeeded/failed page counts and crawledAt. Additional evidence, all-profile arrays and typed repository links support consumers that need more context.
This is not a full drop-in replacement. Keep these differences in your integration:
| Area | Migration consideration |
|---|---|
| Coverage | Static HTML, host scope, robots rules, available links, page limits and filters affect which contacts are returned. No all-sites or permanent success guarantee. |
| Phones/socials | Phones are normalized to E.164. A representative social link may differ; use all-profile/repository arrays and source evidence. Links do not prove ownership. |
| Email screening | Format/MX heuristics and provider flags can differ. SMTP observations are not mailbox or catch-all verification; live SMTP remains unvalidated. |
| Delivery | Completion order is not input order. Failed rows stay in the dataset; all-failed runs fail after saving diagnostics. The cache is run-local. |
| Concurrency/proxy | Requested concurrency up to 20 shares a 20-request ceiling. Live high-volume throughput and paid proxy routing have not been established. |
Run from Python and automate exports
Install apify-client in your client environment. Set APIFY_TOKEN securely in that environment; never put it in the input or source URLs. The example uses Actor ID OOjMOcPufmpGn0fhg. If desired, set APIFY_MAX_TOTAL_CHARGE_USD to your chosen event-cost cap before execution; it is a budget, not a quoted product price.
This client example uses the Actor's default build. For version-specific behavior, pass an available build number using the client's build argument; a documented version does not by itself change the Actor's default.
import jsonimport osfrom decimal import Decimalfrom apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])options = {}if os.environ.get("APIFY_MAX_TOTAL_CHARGE_USD"):options["max_total_charge_usd"] = Decimal(os.environ["APIFY_MAX_TOTAL_CHARGE_USD"])run = client.actor("OOjMOcPufmpGn0fhg").call(run_input={"urls": ["https://www.scrapingbee.com/"],"maxPagesPerSite": 3,"maxConcurrency": 1,"maxWebsitesConcurrency": 1,"verifyEmails": False,"useProxy": False,},timeout_secs=300,**options,)if run is None:raise RuntimeError("No run record returned")# Read persisted rows even when the run failed after an all-failed scan.for row in client.dataset(run["defaultDatasetId"]).iterate_items():print(json.dumps(row, ensure_ascii=False))store = client.key_value_store(run["defaultKeyValueStoreId"])for key in ("OUTPUT", "SOURCE_DIAGNOSTICS"):record = store.get_record(key)print(key, json.dumps(record["value"] if record else None, ensure_ascii=False))print("Run status:", run["status"])
For recurring refreshes, save the input as an Apify task and schedule it. Keep each run ID with your exports, filter or route failed rows separately, and deduplicate downstream across runs if needed. A fresh run performs a fresh acquisition; prior runs do not populate its cache.
FAQ
Why did a successful scan return no contacts? A readable page may publish none, link to contacts beyond your page limit or require JavaScript. Check the attempted/readable counts, sourcePages, contact-page discovery and stop reason before increasing coverage.
Are all extracted addresses contacts for the target company? No. A site can publish hosting-provider, legal or other third-party addresses. Review emailDomainRelation and source context, or select same-site. Domain matching still does not prove ownership.
Why are repeated URLs charged again in paid mode? The unit is a saved readable input row. The run-local cache reduces repeated acquisition, while each delivered readable copy is a separate website event. Remove duplicates from your input if you only need one row.
Can I use SMTP results as a validated mailing list? No. Neither syntax, MX records nor RCPT acceptance proves delivery. SMTP remains experimental and never returns a verified-mailbox claim.
Why can a failed run still have results? Complete failures are useful records. All-failed runs save them before failing; mixed runs can succeed on the platform while OUTPUT.status is partial. Read the dataset and diagnostics instead of relying only on the run status.