Domain to Verified Work Email avatar

Domain to Verified Work Email

Pricing

from $8.50 / 1,000 published person matches

Go to Apify Store
Domain to Verified Work Email

Domain to Verified Work Email

Find the work email a named person publishes on their own company website, confirmed by a same-document schema.org Person assertion. No pattern guessing, no SMTP probing. $0.01 only when a published match is found; no-match, role mailbox, ambiguous and partial outcomes are free.

Pricing

from $8.50 / 1,000 published person matches

Rating

0.0

(0)

Developer

Tim Zinin

Tim Zinin

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

3

Monthly active users

15 days ago

Last modified

Share

Domain to Verified Work Email: Published-Evidence Person Matching for Company Websites

What this Actor does: input, evidence, outcome Get one work email per person when the company's own public website publishes that literal address next to the person's name. Two tiers carry the address: a published person match, where the same page also asserts the exact name-to-address pairing in machine-readable schema.org Person markup, and a published visible match, where the page prints the address in the person's own card without that markup. When neither exists, the run still returns everything useful it saw, free of charge: name-shaped candidate addresses that could not be attributed cleanly, and the company's role mailboxes such as info@ or press@.

This is not a guessing service. It does not synthesize first.last@domain patterns, does not call MX or SMTP servers, does not run a catch-all probe, and does not treat "the mailbox accepted the message" as proof that a specific human owns it. The only billable outcome today is the published person match. A published visible match is delivered under its own event, visible-match-found, which is billed only while the Store pricing lists it with a price; until then it is free. Everything else the run finds is always free: candidates, role mailboxes, no-match, ambiguous, robots-blocked, partial or budget-stopped.

What you get

For each {fullName, domain} pair you submit, the Actor visits only the buyer-selected company's own website, reading at most four fixed pages in a fixed relevance order — /team, /about, /contact, then / — and returns exactly one of four outcomes:

  • A published person match (published_person_match, billed). The literal email address exists as visible text or an explicit mailto: link on an allowed page. The requested full name appears in the same bounded visible context. The local part of the address follows a deterministic personal-name form (first, last, first.last, first+last, first-initial+last, or last+first-initial). The address is not a role mailbox, placeholder, image filename or third-party domain. And the same HTML document also carries valid bounded JSON-LD with an exact schema.org context, @type of Person (or the full schema.org Person URL), whose normalized name and literal email exactly match the requested person and the candidate address.
  • A published visible match (published_visible_match, its own event). Every condition above holds except the schema.org assertion: the company page prints the address inside the requested person's own block, with no second person competing for it. This is what most real company sites publish. The address is returned in email with confidenceBand:"medium", and the row is billed only while the visible-match-found event carries a Store price.
  • A free candidate list (candidate_address_unconfirmed). The site publishes an address shaped like the person's name (jdoe@, jane@) but the page never prints the name next to it, or a second person shares the context. email stays null; the addresses are listed in candidateEmails for a human to review. Nothing is charged.
  • An honest free decision. No literal address at all, only role mailboxes (listed in roleEmails), an address tied to a different person, two candidates that both clear the same bar (an explicit tie, both listed in candidateEmails), a page the runtime could not read, a page robots.txt disallowed, or a budget stop before or after the source read completed. Every free decision still returns a full row with checked pages, unreachable pages, robots exclusions and a partialReason or plain no-match reason — never a bare empty result.

The Actor never claims mailbox ownership, consent to be contacted, present-day deliverability or willingness to receive outreach. Those four claims are outside what a published web page can prove, so every row — paid or free — carries safeToAutomate:false and safeForOutreach:false as non-negotiable machine flags, not just prose caveats. A published person match is evidence to route to a human for a compliance-checked next step, not a green light for automatic email sending.

Who uses it

  • Sales and partnerships teams building an enrichment step for a CRM record where an unverified guessed address already causes bounce-rate and deliverability damage, and who need an auditable "here is the exact page and the exact machine-readable claim" record before a human approves outreach.
  • Recruiting and sourcing teams confirming a public work address for a named candidate before a manual outreach step, without scraping a personal inbox or a paid people-search database — only what the target company itself already publishes.
  • Compliance and data-quality reviewers who need a defensible reason to accept or reject an address already sitting in a CRM or spreadsheet — this Actor's row is the reason, with source URL, bounded-context hash and observation time attached.
  • Developers and workflow builders who want a strict, flat, machine-checkable contract instead of parsing prose — an agent, an n8n node or an MCP client can branch on verificationStatus and found without guessing at free text.
  • Internal automation agents operating under human review — the Actor is explicitly designed so an LLM-driven workflow can consume its output safely: every non-terminal state is machine-typed, and the two safety flags default to false so an agent cannot silently escalate a free or ambiguous row into an outreach action.

This is a one-person, one-domain, on-demand utility. It is not a monitor, not a bulk company-wide crawler, not a mailbox verifier and not a people-search database. It does not accept a list of guessed addresses to "check" — it only accepts a name and a domain and only ever returns an address it found published, never one it invented.

Where the Actor sits in your pipeline: trigger, run, action

How to run

  1. Open the Actor's Input form (Console, API or an MCP client) and supply one to fifty {fullName, domain} objects, each with an optional companyName hint and an optional inputRef you control for CRM correlation.
  2. Optionally cap maxPagesPerPerson (1–4, default 3) and maxConcurrency (1–5, default 1; platform-billed work is always processed sequentially regardless of this setting — it only affects local, non-monetized batching).
  3. Start the run from the Console Start, the Apify API runs endpoint with your token, an Apify Task on a schedule, or an MCP tool call using the same strict input.
  4. Read the default Dataset for one row per submitted person — every row is present even for duplicates, ambiguous pages or source failures, so row count against input count is itself a completeness check.
  5. Read the key-value store record OUTPUT for one run-level summary: input rows, unique people after de-duplication, duplicates skipped, matched count, free-decision count, partial count, budget-stopped count, whether the run stopped early, and why.
  6. Route every published_person_match and published_visible_match row to a human-reviewed next step, and review candidateEmails and roleEmails by hand. Do not wire any outcome directly into an autosend integration; the safety flags exist specifically to prevent that shortcut.

No login, browser session, proxy or CAPTCHA-solving step is ever required to run this Actor, because none of those techniques are part of its source-access design (see Sources and rights, below). A default-input smoke run does not require any buyer secret.

Pricing

$0.01 per published person match, plus a $0.005 Actor start fee. A published visible match is delivered under a separate event, visible-match-found; it is billed at that event's Store price only while the pricing page lists one, and delivered free otherwise. Every other outcome — candidate list, role mailboxes, no-match, unlinked personal address, ambiguous tie, robots-blocked page, unreachable page, partial coverage, or a budget stop before or after a source read — costs nothing beyond the flat start fee. The default Dataset item event itself is priced at exactly zero, so reading rows back never re-bills you. The chargedEvent field on every row names the event it was billed under, or null when it was free.

The only data source is the company's own public website that you selected — there is no third-party data license or purchased list behind this Actor, so there is nothing to resell and no separate data bill on top of the price above.

The billable unit is always a delivered address — never a page read, never a "lookup," never a flat per-input charge. A run of fifty people where three clear the strict bar and seven the visible bar bills for three published person matches ($0.03), plus the seven visible matches only if that event is priced, plus the one start fee — not fifty.

Input contract

people is required: one to fifty objects. Each object requires fullName (a real first-plus-last name, 3–160 characters, must contain at least two whitespace-separated tokens) and domain (a bare public hostname, 4–253 characters — schemes, paths, credentials and IP literals are all rejected by pattern before any network call). companyName is an optional page-evidence hint, never used to invent an address. inputRef is an optional opaque string, up to 200 characters, copied back to the result row exactly as given — including leading, trailing, repeated or tab whitespace — for CRM/workflow correlation; it is validated but never normalized or interpreted.

You cannot supply a path, a full URL or a page list. The runtime alone selects the fixed, relevance-ordered page set — /team, /about, /contact, / — and maxPagesPerPerson only narrows how many of those four the run is willing to read for a given person, not which ones. maxConcurrency only affects local batching; it never parallelizes billed source work.

Unicode names and internationalized domain names are normalized through the current Public Suffix List, including private-boundary entries. Bare public suffixes, URLs, embedded credentials and IP literals are rejected before any source call. Duplicate {normalized full name, registrable domain} pairs are collapsed once, before any page is fetched, so re-submitting the same person against the same company twice in one run never double-charges you.

{
"people": [
{
"fullName": "King Lee",
"domain": "visionwide.co",
"companyName": "Visionwide",
"inputRef": "crm-lead-40217"
},
{
"fullName": "Dana Osei",
"domain": "examplecorp.com",
"companyName": "Example Corp",
"inputRef": "crm-lead-40218"
}
],
"maxPagesPerPerson": 3,
"maxConcurrency": 1
}

The runtime independently re-validates every field and rejects unknown properties and wrong JSON types even where the platform schema would already have caught them, so a caller bypassing the Console form through the raw API gets the same guarantees as one using it.

Example output

Every row shares one flat 34-field schema regardless of outcome, so a database, dataframe or workflow condition never needs to branch on record shape — only on field values. The rows below are illustrative examples built directly from dataset_schema.json and the matching contract, showing the exact field shape and values this Actor emits for each outcome class.

Published person match

{
"recordType": "person_email_decision",
"schemaVersion": "1.1",
"entityId": "5c64f059712997f402463d54",
"inputRef": "crm-lead-40217",
"fullName": "King Lee",
"companyName": "Visionwide",
"domain": "visionwide.co",
"email": "king.lee@visionwide.co",
"verificationStatus": "published_person_match",
"verificationLevel": "published_person_to_address_match",
"matchScore": 100,
"matchReasons": [
"literal_email_published",
"same_registrable_company_domain",
"full_name_on_page",
"personal_local_part_matches_name",
"schema_org_person_name_email_assertion",
"full_name_in_same_bounded_dom_context",
"company_name_on_page"
],
"sourceUrl": "https://visionwide.co/team",
"sourcePath": "/team",
"sourceEvidence": {
"literalEmailPublished": true,
"fullNameOnPage": true,
"personalLocalPart": true,
"publicationChannel": "visible_text",
"boundedContextSha256": "8b60d44f3fd265e726822263b87b1cf471d25ff506ed91690068718154d27246",
"fullNameInSameContext": true,
"schemaOrgPersonAssertion": true,
"disclaimer": "Published-page match with a same-document schema.org Person assertion; not proof of mailbox ownership, consent, or deliverability."
},
"candidateEmails": [],
"roleEmails": ["info@visionwide.co"],
"checkedPages": [
{ "path": "/team", "url": "https://visionwide.co/team", "status": 200, "truncated": false },
{ "path": "/about", "url": "https://visionwide.co/about", "status": 200, "truncated": false }
],
"unreachablePages": [],
"robotsDisallowed": [],
"confidenceScore": 100,
"confidenceBand": "high",
"recommendedAction": "ingest_evidence_then_apply_your_legal_outreach_policy",
"safeToAutomate": false,
"safeToIngestEvidence": true,
"safeForOutreach": false,
"partial": false,
"partialReason": null,
"errorCode": null,
"error": null,
"retryable": false,
"observedAt": "2026-08-11T14:32:07.418Z",
"found": true,
"chargedEvent": "result-found"
}

Every key, value shape and const in this row matches the actual dataset_schema.json conditional rules for published_person_match — the exact seven matchReasons tokens, the exact eight-key sourceEvidence object, the fixed recommendedAction string, and matchScore/confidenceScore/ confidenceBand pinned to 100/100/"high" are the only values the schema allows for this row. chargedEvent is "result-found" on the platform and null in a local, non-monetized run.

Published visible match

{
"recordType": "person_email_decision",
"schemaVersion": "1.1",
"entityId": "9a1c3f5e7b2d4c6a8e0f1b3d",
"inputRef": "crm-lead-40219",
"fullName": "Colin Percival",
"companyName": "Tarsnap",
"domain": "tarsnap.com",
"email": "cperciva@tarsnap.com",
"verificationStatus": "published_visible_match",
"verificationLevel": "published_visible_name_to_address_match",
"matchScore": 90,
"matchReasons": [
"literal_email_published",
"same_registrable_company_domain",
"full_name_on_page",
"personal_local_part_matches_name",
"full_name_in_same_bounded_dom_context",
"company_name_on_page"
],
"sourceUrl": "https://www.tarsnap.com/contact",
"sourcePath": "/contact",
"sourceEvidence": {
"literalEmailPublished": true,
"fullNameOnPage": true,
"personalLocalPart": true,
"publicationChannel": "mailto",
"boundedContextSha256": "2f7d1c9e4b8a6053e1d2c3b4a5968778695a4b3c2d1e0f9a8b7c6d5e4f3a2b1c",
"fullNameInSameContext": true,
"schemaOrgPersonAssertion": false,
"disclaimer": "Published-page match: the company site prints this address next to the requested name, without a schema.org Person assertion; not proof of mailbox ownership, consent, or deliverability."
},
"candidateEmails": [],
"roleEmails": [],
"checkedPages": [
{ "path": "/team", "url": "https://www.tarsnap.com/team", "status": 404, "truncated": false },
{ "path": "/about", "url": "https://www.tarsnap.com/about", "status": 404, "truncated": false },
{ "path": "/contact", "url": "https://www.tarsnap.com/contact", "status": 200, "truncated": false }
],
"unreachablePages": [
{ "path": "/team", "status": 404, "errorCode": "HTTP_404", "error": "Source returned HTTP 404", "retryable": false },
{ "path": "/about", "status": 404, "errorCode": "HTTP_404", "error": "Source returned HTTP 404", "retryable": false }
],
"robotsDisallowed": [],
"confidenceScore": 90,
"confidenceBand": "medium",
"recommendedAction": "review_evidence_then_apply_your_legal_outreach_policy",
"safeToAutomate": false,
"safeToIngestEvidence": true,
"safeForOutreach": false,
"partial": false,
"partialReason": null,
"errorCode": null,
"error": null,
"retryable": false,
"observedAt": "2026-09-15T20:10:41.002Z",
"found": true,
"chargedEvent": null
}

The visible tier carries the same evidence object as the strict tier with schemaOrgPersonAssertion:false, scores 85 (or 90 when the page also names the company), lands in confidenceBand:"medium", and is billed under visible-match-found only while that event is priced — here chargedEvent is null, so the buyer paid nothing for the row. Missing pages that returned 404 are recorded in both checkedPages and unreachablePages and do not make the row partial.

Free — candidate addresses listed for review

{
"recordType": "person_email_decision",
"schemaVersion": "1.1",
"entityId": "4b8e2d6f0a1c3e5b7d9f1a2c",
"inputRef": "crm-lead-40220",
"fullName": "Mara Okafor",
"companyName": "Northwind Labs",
"domain": "northwindlabs.io",
"email": null,
"verificationStatus": "candidate_address_unconfirmed",
"verificationLevel": "none",
"matchScore": 0,
"matchReasons": [],
"sourceUrl": null,
"sourcePath": null,
"sourceEvidence": {
"candidateCount": 0,
"candidateAddressCount": 1,
"roleAddressCount": 2,
"unlinkedPersonalAddressCount": 0,
"thirdPartyAddressCount": 0,
"disclaimer": "Name-shaped addresses were published on the company site, but none is attributed to the requested person clearly enough to return as a match; review them manually."
},
"candidateEmails": ["mokafor@northwindlabs.io"],
"roleEmails": ["hello@northwindlabs.io", "press@northwindlabs.io"],
"checkedPages": [
{ "path": "/team", "url": "https://northwindlabs.io/team", "status": 200, "truncated": false },
{ "path": "/about", "url": "https://northwindlabs.io/about", "status": 200, "truncated": false },
{ "path": "/contact", "url": "https://northwindlabs.io/contact", "status": 200, "truncated": false }
],
"unreachablePages": [],
"robotsDisallowed": [],
"confidenceScore": 0,
"confidenceBand": "none",
"recommendedAction": "review_candidate_addresses_manually",
"safeToAutomate": false,
"safeToIngestEvidence": false,
"safeForOutreach": false,
"partial": false,
"partialReason": null,
"errorCode": null,
"error": null,
"retryable": false,
"observedAt": "2026-09-15T20:12:03.517Z",
"found": false,
"chargedEvent": null
}

Here the /team page names Mara Okafor, but mokafor@ is printed in a footer far from her card, so it is a candidate rather than a match. The row is free, email stays null, and both the candidate and the two role mailboxes are handed to a human reviewer instead of being discarded.

Partial — coverage could not be completed

{
"recordType": "person_email_decision",
"schemaVersion": "1.1",
"entityId": "2ec54ae61eb6352a443897e7",
"inputRef": "crm-lead-88031",
"fullName": "Priya Chandran",
"companyName": null,
"domain": "northfieldstudio.io",
"email": null,
"verificationStatus": "partial_source_coverage",
"verificationLevel": "none",
"matchScore": 0,
"matchReasons": [],
"sourceUrl": null,
"sourcePath": null,
"sourceEvidence": {
"candidateCount": 0,
"candidateAddressCount": 0,
"roleAddressCount": 0,
"unlinkedPersonalAddressCount": 0,
"thirdPartyAddressCount": 0,
"disclaimer": "One allowed page could not be read before the page budget was exhausted; no candidate is eligible for payment from an incomplete scan."
},
"candidateEmails": [],
"roleEmails": [],
"checkedPages": [
{ "path": "/team", "url": "https://northfieldstudio.io/team", "status": 200, "truncated": false },
{ "path": "/contact", "url": "https://northfieldstudio.io/contact", "status": null, "truncated": false }
],
"unreachablePages": [
{
"path": "/contact",
"status": null,
"errorCode": "SOURCE_TIMEOUT",
"error": "Source request timed out",
"retryable": true
}
],
"robotsDisallowed": [],
"confidenceScore": 0,
"confidenceBand": "none",
"recommendedAction": "retry_source_later",
"safeToAutomate": false,
"safeToIngestEvidence": false,
"safeForOutreach": false,
"partial": true,
"partialReason": "source_page_unavailable",
"errorCode": "SOURCE_TIMEOUT",
"error": "One allowed page (/contact) timed out before the run established complete coverage; the pages that were read are still reported above.",
"retryable": true,
"observedAt": "2026-08-11T14:41:52.005Z",
"found": false,
"chargedEvent": null
}

Notice that /contact appears in both checkedPages (attempted, no status obtained) and unreachablePages (the failure detail) — the schema requires both entries together so a consumer never has to infer an attempt from a failure record alone. errorCode/error/retryable here use the exact bounded SOURCE_TIMEOUT family the schema defines for a genuine timeout, not a free-text guess.

Free — role mailbox found, no person match

{
"recordType": "person_email_decision",
"schemaVersion": "1.1",
"entityId": "f31dd870b14954a167a92211",
"inputRef": "crm-lead-40218",
"fullName": "Dana Osei",
"companyName": "Example Corp",
"domain": "examplecorp.com",
"email": null,
"verificationStatus": "role_address_only",
"verificationLevel": "none",
"matchScore": 0,
"matchReasons": [],
"sourceUrl": null,
"sourcePath": null,
"sourceEvidence": {
"candidateCount": 0,
"candidateAddressCount": 0,
"roleAddressCount": 1,
"unlinkedPersonalAddressCount": 0,
"thirdPartyAddressCount": 0,
"disclaimer": "No email is returned without one unique evidence-backed person match."
},
"candidateEmails": [],
"roleEmails": ["info@examplecorp.com"],
"checkedPages": [
{ "path": "/team", "url": "https://examplecorp.com/team", "status": 200, "truncated": false },
{ "path": "/about", "url": "https://examplecorp.com/about", "status": 200, "truncated": false },
{ "path": "/contact", "url": "https://examplecorp.com/contact", "status": 200, "truncated": false }
],
"unreachablePages": [],
"robotsDisallowed": [],
"confidenceScore": 0,
"confidenceBand": "none",
"recommendedAction": "review_or_supply_a_different_public_source",
"safeToAutomate": false,
"safeToIngestEvidence": false,
"safeForOutreach": false,
"partial": false,
"partialReason": null,
"errorCode": null,
"error": null,
"retryable": false,
"observedAt": "2026-08-11T14:44:19.771Z",
"found": false,
"chargedEvent": null
}

Notice what stays constant across all five rows: the same 34 keys, the same nullable-not-missing shape, and the same pair of safety flags fixed at false. A consumer never has to sniff which fields exist before reading a row. The role mailbox is no longer discarded: it is listed in roleEmails so the buyer leaves with a usable company contact even when no person-level address exists.

Field dictionary

FieldMeaningImportant boundary
recordTypeAlways "person_email_decision"Constant discriminator; safe for schema routing
schemaVersionAlways "1.1"Bumped from 1.0 when the visible tier, candidate list and role list were added
entityIdStable 24-hex-character row identityDeduplication key, not a CRM record ID
inputRefYour caller reference, copied exactlyNever normalized; whitespace preserved on purpose
fullNameNormalized requested nameEcho of input, not a claim the person exists at the domain
companyNameOptional company hint you suppliedEvidence signal only; never used to fabricate an address
domainNormalized registrable company domainPSL-normalized; IDN and Unicode handled before matching
emailThe literal published addressPresent only on published_person_match and published_visible_match; null on every other row — never a guessed pattern
verificationStatusOne of twelve deterministic outcome codespublished_person_match bills result-found; published_visible_match bills visible-match-found only while priced
verificationLevelpublished_person_to_address_match, published_visible_name_to_address_match or noneDescribes what kind of evidence was found, not a confidence score
matchScore0–100 deterministic evidence scoreStrict rows score exactly 100; visible rows 85, or 90 when the page also names the company
matchReasonsNamed evidence tags for a matchEmpty array on every free row
sourceUrlThe exact page the match came fromnull unless one page carries the full assertion
sourcePathWhich of the four allowed paths matchednull, /team, /about, /contact or / — never a caller-supplied path
sourceEvidenceStructured evidence object with a mandatory disclaimerdisclaimer is present on every row, paid or free
checkedPagesOrdered list of pages actually readUnique paths only, drawn from the fixed four-path sequence
unreachablePagesPages that could not be read, with error detailEvery entry's errorCode/error/retryable triple is drawn from one closed, schema-enforced family — never a free-text guess
robotsDisallowedAllowed-set paths robots.txt blockedThe Actor never reads a disallowed path to "check anyway"
confidenceScore0–100 routing-confidence numberNot a probability-of-correctness statistic; a deterministic bucket key
confidenceBandhigh, medium, ambiguous or nonehigh = strict tier, medium = visible tier
recommendedActionOne of a small closed set of queue labels, pinned by verificationStatus/retryableA routing tag for a human queue, never an authorization to act — the exact string per status is schema-enforced, not freely chosen at write time
safeToAutomateAlways falseFixed by the Dataset schema; no future row can flip this without a schema version bump
safeToIngestEvidenceWhether the row's evidence is complete enough to store downstreamDistinct from safeForOutreach — evidence can be ingestible without being outreach-ready
safeForOutreachAlways falseThis product never certifies a message is safe to send
partialWhether complete allowed-page coverage was establishedA true row can never carry verificationStatus:"published_person_match"
partialReasonWhy coverage was incompletenull unless partial:true; otherwise exactly one of source_page_unavailable, source_page_truncated, unsupported_content_type, robots_disallowed_page, robots_unavailable_fail_closed, matcher_work_limit_exceeded, budget_stopped_before_source_read, budget_stopped_after_source_read, person_processing_error
errorCodeBounded machine error codenull on a clean decision; present on every unreachable/failure row
errorShort bounded human-readable error textNever contains request headers, cookies or page bodies
retryableWhether a bounded retry is appropriateDoes not override the run's own budget and deadline caps
observedAtWhen this Actor observed the sourceDistinct from any timestamp the source page itself displays
foundWhether the row returned an address in emailtrue on both address-bearing tiers; gate payment logic on chargedEvent, not on this flag
candidateEmailsUp to five name-shaped addresses on the company domain that could not be attributed cleanlyAlways free; empty on address-bearing rows; two or more on an explicit tie
roleEmailsUp to five role mailboxes (info@, press@, …) found on the company's own domainAlways free; a human fallback, never a person-level claim
chargedEvent"result-found", "visible-match-found" or nullThe event this exact row was billed under; null means the row cost nothing

Evidence and boundaries

What "verified" does and does not mean here. In this product, verified means visible published person-to-address evidence on the company's own site, with source URL and observation time — strengthened, in the strict tier, by an exact schema.org Person name-email assertion on the same page. It does not mean mailbox ownership, consent, present deliverability or willingness to receive outreach. That sentence is not marketing language — it is copied from the product contract and repeated in every accepted row's disclaimer field so the boundary travels with the data, not just this page.

Why there are two address-bearing tiers. A literal address next to a name in ordinary prose is real evidence, but weaker than a machine-readable claim: company pages sometimes list people near addresses that belong to someone else, near role mailboxes, or near stale contact blocks. The strict tier therefore requires the same page to independently assert the exact pairing through schema.org/Person JSON-LD, and only that tier is billed as a published person match. The visible tier applies every other check — literal address, personal-name local part, the name in the same bounded block, no competing person in that block, no conflicting structured claim elsewhere — and returns the address with confidenceBand:"medium". Because most company sites publish contacts without JSON-LD, the visible tier is where the majority of real answers come from; the strict tier is where the strongest ones come from. If a page has the name, has the address, and even has a JSON-LD block that does not exactly match, the row is a visible match, not a strict one.

Why candidates are listed instead of discarded. An address shaped like the requested person (jdoe@, jane@) that the page never prints next to the name, or prints in a block another person shares, is not attributed well enough to return as email. Earlier versions counted such addresses and threw them away. They are now listed in candidateEmails, free, so a reviewer can decide; the Actor still never promotes them to a match on its own.

Why hidden content is excluded before matching. Real company pages hide old staff blocks, template placeholders and legacy contact cards with display:none, zero-opacity layers, off-screen positioning, closed <details> elements and CSS techniques ranging from simple inline styles to !important cascades and CSS-escape obfuscation. Inline styles use the bounded deterministic matcher. When a page links stylesheets, the Actor fetches up to 96 public HTTPS sheets through per-host DNS, SSRF, robots and byte guards, then applies them in an offline Chromium page with JavaScript and every browser network request disabled. Elements whose computed display, visibility, content-visibility or opacity hides them are removed before matching. A sheet that cannot be fetched under those guards (a robots-blocked CDN, an @import, an oversized file) is skipped rather than failing the page; a page with skipped sheets can still yield a visible match, but never the strict, schema.org-billed tier, because the skipped CSS might have hidden a stale block. Earlier versions failed every page with more than 32 sheets or one blocked sheet as a free partial result, which made most WordPress sites unreadable.

Why role addresses and weak local parts are rejected. contact-us, customer-care, editorial, recruiters, human-resources, admissions and a defined family of similar role-token variants never become paid rows, even if a name happens to sit nearby. Weak local-part forms (first-name-only or last-name-only) additionally require the requested full name in the immediate bounded evidence context and fail closed if that context contains a conflicting person name. This is a deliberate asymmetry: the bar for treating something as a real personal address is higher than the bar for treating it as merely present text.

Why the source path is narrow by design. The Actor reads only the buyer-selected company's own public website, bounded to the same registrable domain, at most four fixed pages (/team, /about, /contact, /), over HTTPS on the default port only. Every hostname is resolved once per run and pinned to globally routable addresses before any connection; private, loopback, metadata, documentation, transition and mixed private/public resolutions are all blocked before a socket opens. Redirects are followed manually, only within the same public-suffix registrable-domain boundary, and every hop re-verifies scheme, domain and address. There is no browser navigation, no JavaScript execution, no login, no CAPTCHA-solving and no proxy. Chromium is used only as an offline CSS layout engine after guarded HTTP reads, and all of its network requests are aborted. The run retains a hard cap of 300 actual HTTP requests.

Why robots.txt is honored strictly, not politely. The policy file is fetched fail-closed: if it is ambiguous, wrongly typed, redirected through a login page, unreachable, or disallows the target path, that path is simply not read — not read-with-a-warning. A site that simply has no robots.txt (a 404 or 410 at the canonical path) places no restrictions under RFC 9309, and is read normally; earlier versions treated that as a blocked source and returned nothing. The compiled policy respects exact user-agent groups over wildcard groups, closes ambiguous or malformed directive groups before any page request, and matches wildcard/end-anchor rules with literal-segment matching rather than compiling caller-influenced text as a regular expression.

Why two identical evidence rows are not each other's proof of correctness. If two candidate addresses on the same domain both clear the same tier for the requested name, the run does not guess a winner — it returns an explicit free tie and lists both in candidateEmails. Certainty, not volume of near-misses, is what this product sells.

Decision routing

Every row's verificationStatus maps to exactly one of the following twelve outcomes. A downstream workflow should switch on this field directly rather than re-deriving it from email, found or prose.

verificationStatusBillable?What it meansrecommendedAction you will see
published_person_matchYes (result-found)Complete, unambiguous, schema.org-backed evidenceingest_evidence_then_apply_your_legal_outreach_policy
published_visible_matchOnly while visible-match-found is pricedAddress printed in the person's own block, no schema.org assertionreview_evidence_then_apply_your_legal_outreach_policy
candidate_address_unconfirmedNoName-shaped addresses exist but none is attributed cleanly; see candidateEmailsreview_candidate_addresses_manually
ambiguous_multiple_matchesNoTwo or more candidates cleared the same tier; both listed in candidateEmailsretry_source_later if retryable:true, else review_or_supply_a_different_public_source
partial_match_not_billableNoA match exists but coverage was incomplete; the address is listed in candidateEmailsretry_source_later if retryable:true, else review_or_supply_a_different_public_source
partial_source_coverageNoNot every allowed page could be readretry_source_later if retryable:true, else review_or_supply_a_different_public_source
role_address_onlyNoOnly role mailboxes were found; see roleEmailsreview_or_supply_a_different_public_source
unlinked_personal_address_onlyNoPersonal-shaped addresses exist but none is shaped like the requested namereview_or_supply_a_different_public_source
no_published_person_matchNoComplete coverage, genuinely nothing foundreview_or_supply_a_different_public_source
budget_stopped_before_source_readNoThe run's cap was reached before this person was processedraise_max_total_charge_or_reduce_input
budget_stopped_after_source_readNoThe cap was reached after reading pages but before a result could be deliveredraise_max_total_charge_or_reduce_input
partial_processing_errorNoAn internal processing limit or error stopped this person's evaluationretry_source_later

Every value in that last column is a schema-enforced constant, not free text — the Dataset schema pins the exact string per status (and, for the three retry-eligible statuses, per retryable value too), so a downstream switch statement can match on it exactly rather than pattern-matching prose.

A workflow implementation only needs a handful of rules to stay correct:

  • Only continue an automated pipeline step on found === true and partial === false, and branch on confidenceBand when the strict and visible tiers should take different paths.
  • A non-null email appears only on published_person_match and published_visible_match under the enforced schema; treat any other combination as a hard integration bug, not a valid state.
  • Gate spend reconciliation on chargedEvent, never on found: a visible match can be found:true and still chargedEvent:null.
  • Treat candidate_address_unconfirmed, ambiguous_multiple_matches and unlinked_personal_address_only as "evidence exists but is not sufficient," not as a weaker version of a match — they are unbillable, and none is eligible for an autosend tier.
  • Retry only rows where retryable === true, and only after the specific delay/backoff your workflow already applies to other bounded-source jobs — this Actor's own run-level cap does not retry automatically.
  • Deduplicate downstream on entityId, not on {fullName, domain}, since the entity ID is stable across identical resubmission while your own upstream normalization might differ slightly.

Commercial playbooks

CRM email-quality gate

Before a guessed or purchased address is allowed into an active outreach sequence, run it — or the name/domain pair it came from — through this Actor. A published_person_match becomes supporting evidence for a human reviewer's decision to proceed; anything else becomes a reason to hold the record for manual verification rather than let a guessed pattern degrade sender reputation. This playbook explicitly does not wire the result into an autosend step; the Actor's own safety flags are designed to stop that shortcut, and the human review step is not optional.

Recruiting candidate-contact confirmation

When a sourcing workflow already has a candidate's public profile and target company, this Actor checks whether that same company's own site independently publishes a matching, schema-asserted address for that exact name. A match gives a recruiter a defensible, source-linked reason to reach out through the confirmed channel instead of a purchased database guess; a free result is a signal to fall back to the platform's native messaging instead of fabricating an email.

Partner and vendor identity verification

Before onboarding a named contact from a partner or vendor company into a billing, support or security-notification list, this Actor confirms that the company's own website — not a third-party aggregator — currently publishes that person at that address with a machine-readable assertion. This is a lightweight identity-hygiene check, not a KYC or sanctions-screening substitute.

CRM data-quality audit sweep

Run a batch of existing CRM contacts (name + company domain, using inputRef to carry the CRM record ID) to find stale or fabricated-looking addresses. Rows with no_published_person_match or role_address_only for a contact your CRM currently lists as a personal, verified address are a worklist for data-hygiene review — not an automatic field overwrite, since the source page not currently publishing an address does not prove the address itself is wrong.

Support and escalation contact confirmation

For B2B accounts where an internal contact list has gone stale, this Actor can re-confirm whether a named account contact still has a personal address published on the vendor's own site before an account manager escalates an issue directly instead of through a shared support alias.

Integration recipes

The Actor is built as a machine-first microservice: a strict, closed Input schema; a uniform, strictly typed Dataset; and a separately schema-validated OUTPUT summary record. Callers can drive it through the Apify Console, the Apify API, a scheduled Apify Task, or an MCP client using the same Input contract described above.

Scheduled CRM enrichment pattern

  1. A buyer-owned Apify Task stores the strict Input — no secrets are required by this Actor's own contract, since it reads only public pages.
  2. A schedule starts a bounded run against a queue of {fullName, domain, inputRef} triples pulled from CRM records that currently lack a verified address.
  3. The workflow waits for terminal success and reads KVS OUTPUT for matched, freeDecisions, partial and budgetStopped counts before touching Dataset rows.
  4. It reads Dataset rows and routes every published_person_match into a human review queue keyed by inputRef.
  5. It logs every free verificationStatus value back to the CRM record as a "checked, not verified" marker rather than leaving the field silently untouched.

Event-driven single-lookup pattern

  1. A new-lead webhook or form submission triggers a single-person run with one people entry.
  2. The caller polls run status or awaits a webhook/finish callback, then reads the one Dataset row.
  3. If found === true, the lead record gets a "company-verified address available for review" flag; the address itself is still routed to a human before it enters an outreach sequence.
  4. If partial === true and retryable === true, the caller requeues once after its own standard backoff window rather than treating the gap as a permanent no-match.

Agent/MCP pattern

  1. The agent supplies exactly the fields the Input schema allows — it cannot inject a path, a credential or an alternate protocol even if it tries, because the schema and the runtime validator both reject anything outside the closed contract.
  2. It reads found, partial, verificationStatus, safeToAutomate and safeForOutreach before deciding on any next step.
  3. It never converts a role_address_only or unlinked_personal_address_only row into an affirmative claim just because an address string exists in the row's evidence object.
  4. It always routes a published_person_match to a human queue rather than an autonomous send action, honoring safeForOutreach:false as a hard stop, not a suggestion.
  5. It cites entityId, sourceUrl and observedAt in any work item or ticket it creates, so a human reviewer can re-check the original page directly.

Data-warehouse pattern

Append Dataset rows keyed by entityId. Keep observedAt (this Actor's observation time) distinct from any timestamp your own pipeline assigns on ingestion. Preserve null fields as null rather than coercing them to empty strings, since a null email is a structurally different fact from an empty string one. Treat a future schemaVersion bump as an explicit migration, not a silently compatible row shape.

Operating guide

Choosing maxPagesPerPerson. The default of 3 reads the three highest-relevance pages before the homepage fallback. Raising it to 4 adds the homepage as a fourth attempt for companies whose staff information lives only there; lowering it to 1 or 2 trades completeness for a tighter page budget when you already have strong reason to believe the address lives on /team specifically.

Reading the OUTPUT summary before Dataset rows. The run-level summary tells you, in one object, whether the whole batch behaved as expected — inputRows, uniquePeople, duplicateInputsSkipped, matched (strict), visibleMatched, candidateDecisions, billedRows, freeDecisions, partial, budgetStopped, stopped and stopReason — before you spend time iterating individual rows. A stopped:true run with a non-null stopReason means the batch did not fully complete and some people at the tail of your input list may never have been processed; check inputRows against the actual Dataset row count to confirm.

Understanding replay safety. Every paid delivery is journaled in two phases before it ever reaches you: source_reserved before any page is read, then delivery_pending with the exact row hash before the atomic Dataset write, and only delivered after a confirmed platform receipt. If a worker process is interrupted between the atomic write and receiving its receipt, the next process reconciles a bounded read of the Dataset before doing any new source work, recognizes a row whose hash matches the pending entry as already delivered, and skips it rather than re-billing you. An unresolved mismatch between the journal and the Dataset becomes a fatal, non-retried state rather than a silent duplicate charge — this Actor would rather stop than guess.

Understanding the budget-stop states. If your run's result cap is reached before a person is processed at all, that row is budget_stopped_before_source_read and no page was ever fetched for it. If the cap is reached after pages were already read for that person but before a decision could be delivered, the row is budget_stopped_after_source_read and preserves every checked/unreachable page fact it gathered — the evidence of what was actually read is never discarded just because the budget ran out at the wrong moment.

Understanding concurrency. maxConcurrency only governs local, non-monetized batching of work before it reaches the platform boundary. On-platform source reads and paid deliveries are always sequential: another source read only starts while at least one more possible result still fits under your run's cap, so you cannot be charged for parallel work that raced past your intended limit.

Interpreting checkedPages and unreachablePages together. These two arrays are drawn from the same fixed four-path universe and are designed to be read together: every checked page that returned a non-200 status has a matching unreachablePages entry, and every unreachablePages entry that carries a status has a matching checkedPages witness. A path never appears as both fully checked with a 200 status and unreachable — treat any apparent contradiction in your own tooling as a bug in how you're reading the row, not an ambiguity the Actor itself would ever produce.

FAQ

Does this Actor guess email address patterns?

No. It never synthesizes first.last@domain or any other pattern. Every address it returns is one it found as literal published text or an explicit mailto: link on the company's own site.

No, and this is the single most important boundary of the product. "Verified" here means published, page-linked, schema.org-asserted evidence — not that the mailbox exists in a working state, not that the named person controls it, and not that they have agreed to receive anything. safeForOutreach is fixed to false on every row for exactly this reason.

Can I automatically send outreach based on a published_person_match row?

The Dataset schema is deliberately built to make that hard to do by accident. safeToAutomate and safeForOutreach are both fixed false, and recommendedAction labels are queue tags for a human, not authorization strings for an automated send integration. Treat every paid row as "ready for a human to review," not "ready to message."

What happens if a person has two published addresses that both look valid?

The run returns an explicit ambiguous_multiple_matches free decision rather than guessing between them, and lists both addresses in candidateEmails. Certainty is a condition of billing, not a tiebreaker applied after the fact.

Does it read pages I don't ask for?

No. The runtime alone selects the exact fixed set — /team, /about, /contact, / — in that relevance order, bounded by maxPagesPerPerson. You cannot supply an arbitrary path or URL, and the Actor never expands its reach beyond the same registrable domain even across redirects.

What happens if robots.txt disallows a page?

That page is never read, and it appears in robotsDisallowed rather than checkedPages. This applies even if a published address would very likely exist there — the source policy is honored strictly, not selectively. A site with no robots.txt at all is unrestricted and is read normally.

Why does a "no match" row still cost the start fee?

The $0.005 start fee covers the run/session overhead itself, independent of outcome; the $0.01 per-person charge is scoped to the published person match, and the visible-match event to the visible tier. Every free decision still costs the flat start fee and nothing per result — and it now carries the candidate and role addresses the run saw, so a "no match" row is rarely empty.

Is browser automation, a proxy or a login step ever used?

It never automates a remote page. Guarded HTTPS reads fetch the bounded HTML and permitted CSS; offline Chromium evaluates CSS with JavaScript disabled and all browser network requests aborted. There is no login, CAPTCHA-solving or proxy configuration surface in the Input schema.

What is entityId and can I use it as a permanent database key?

It is a stable 24-character hex identity generated for the row within this Actor's own delivery and replay-safety journal. It is safe to use for deduplication and idempotency inside your own pipeline; it is not a CRM contact ID and carries no meaning outside this Actor's context.

How is this different from a plain "find contact page and extract emails" scraper?

A plain extractor bills for any literal address it can find anywhere on a page. This Actor attributes each address to the requested person with a bounded-context check (the name must be in the address's own block, with no competing person), separates that visible tier from the stricter schema.org-asserted tier, labels every weaker address as a free candidate instead of a result, and never charges for an address it could not attribute. The strict tier is the narrower, more defensible claim; the visible tier is the one most real company sites can actually satisfy.

Can an AI agent safely call this Actor unattended?

Yes, for the lookup step itself — the strict input, flat output and conservative safety flags are designed for exactly that. The one hard rule an agent must keep is routing published_person_match and published_visible_match results to a human before any outreach action, because the Actor's own flags will not silently authorize that step for it.

Sources and rights

The only data source is the buyer-selected company's own public website, bounded to the same registrable domain, over HTTPS on the default port, limited to four fixed relevance-ordered paths per person. robots.txt is honored fail-closed on every hostname before any page is read. No private, internal, metadata or non-public network destination can ever be reached — every hostname is resolved once and every resolved address is checked against a blocklist covering private, loopback, metadata, documentation, transition and deprecated/retired address ranges before a connection is attempted, and every redirect hop repeats that check rather than trusting the first one. The Actor does not use a proxy, login flow or CAPTCHA-solving step, and it never reads a page path outside the fixed four-path set regardless of what a caller supplies. Its browser engine is an offline CSS evaluator only and cannot initiate source requests.

The Dataset stores a SHA-256 hash of the bounded evidence context rather than the surrounding page text, and it never stores unrelated data incidentally observed on the page — no phone numbers, no addresses of other people, no page content beyond what the matching contract itself requires as evidence. The only addresses a row lists besides email are addresses on the buyer-selected company's own domain that are either shaped like the requested person's name (candidateEmails) or role mailboxes of that company (roleEmails), capped at five each. Buyers are responsible for using the returned evidence in a manner consistent with the target company's own terms of use and applicable data-protection law in their jurisdiction; this Actor's narrow read scope and strict robots.txt compliance reduce but do not eliminate that responsibility, and the human-review requirement built into every safety flag exists specifically to keep a person in the loop before any contact action is taken.

Limits

  • One to fifty people per run; one company domain per person, never a company-wide crawl.
  • At most four fixed, Actor-selected pages per person — never a caller-supplied path.
  • Static HTML and guarded CSS only; offline Chromium runs without JavaScript or network access.
  • A single run-wide cap of 300 actual HTTP requests (250 logical robots/page reads before redirects).
  • No pattern-guessed addresses, no MX/SMTP/catch-all probing, no deliverability claim of any kind.
  • At most five candidate addresses and five role mailboxes listed per row, all on the company's own domain.
  • No claim of mailbox ownership, consent or outreach permission — ever, on any row.
  • safeToAutomate and safeForOutreach are fixed false for every row the Dataset schema allows.
  • No bulk company-wide extraction mode; this is a per-person, per-domain utility, not a monitor.

Support boundary

This Actor is independently reviewed before release. Support covers deterministic input validation, the source-access and matching-contract behavior described in this page, Dataset/Output/KVS schemas, and delivery/replay evidence. It cannot decide whether a specific outreach action is appropriate for your jurisdiction or industry, cannot restore access to a target company's website if it changes its own robots.txt policy or page structure, cannot guarantee that a company continues to publish an address it published at observation time, and cannot provide legal advice on contact or data-protection obligations. When reporting a problem, include the Actor run ID, the redacted entityId, verificationStatus, errorCode and approximate observedAt time. Never send full page HTML, unredacted addresses beyond what you already hold, or any credential — this Actor's own contract never requires one, and support will never ask for one either.