Website Contact Scraper - Email, Phone & Social Media Extractor avatar

Website Contact Scraper - Email, Phone & Social Media Extractor

Pricing

$10.00 / 1,000 results

Go to Apify Store
Website Contact Scraper - Email, Phone & Social Media Extractor

Website Contact Scraper - Email, Phone & Social Media Extractor

Extract contact data from any website: emails, phone numbers and social profiles. One row per domain, each email typed and traced to the page it came from, phones in E.164, socials sorted by network. Crawls up to 15 pages per site, starting with contact, about, legal and privacy pages.

Pricing

$10.00 / 1,000 results

Rating

0.0

(0)

Developer

My Smart Digital

My Smart Digital

Maintained by Community

Actor stats

2

Bookmarked

25

Total users

5

Monthly active users

2 days ago

Last modified

Share

Website Contact Scraper – Email, Phone & Social Media Extractor

Give this Actor a list of website URLs. For each one you get a single row holding the emails, phone numbers and social media profiles published on that site — every value carrying the exact page it was found on.

Feed it the domains you already have and export one CSV where one row is one company: ready for a CRM import, a prospecting list, or an audit of what your own sites expose publicly.

No browser, no proxy, no configuration: a site is read in about 1 to 5 seconds.

One row, for real

Below is an actual result for https://www.mozilla.org, shortened to keep it readable — 15 pages read in 3.2 seconds. Nothing here is an illustration: these are the fields and the values the Actor returned.

{
"domain": "mozilla.org",
"finalUrl": "https://www.mozilla.org/en-US/",
"keyPages": {
"contact": "https://www.mozilla.org/en-US/contact/",
"about": "https://www.mozilla.org/en-US/about/",
"legal": "https://www.mozilla.org/en-US/about/legal/",
"privacy": "https://www.mozilla.org/en-US/privacy/websites/cookie-settings/"
},
"pagesVisited": [
"https://www.mozilla.org/en-US/",
"https://www.mozilla.org/en-US/contact/",
"https://www.mozilla.org/en-US/about/legal/report-infringement/"
],
"emails": [
{
"value": "licensing@mozilla.org",
"type": "general",
"priority": "primary",
"signals": ["text", "same_domain"],
"sourceUrl": "https://www.mozilla.org/en-US/about/legal/terms/mozilla",
"snippet": "tions on Mozilla authored content can be sent to: licensing@mozilla.org.",
"foundIn": "text"
},
{
"value": "dmcanotice@mozilla.com",
"type": "general",
"priority": "secondary",
"signals": ["mailto"],
"sourceUrl": "https://www.mozilla.org/en-US/about/legal/report-infringement",
"snippet": "If you prefer, you can email dmcanotice@mozilla.com.",
"foundIn": "mailto"
}
],
"primaryEmail": "licensing@mozilla.org",
"phones": [
{
"valueRaw": "+1 650-903-0800",
"valueE164": "+16509030800",
"priority": "primary",
"signals": ["text"],
"sourceUrl": "https://www.mozilla.org/en-US/about/legal/report-infringement",
"snippet": "Mozilla's Designated Agent's phone is +1 650-903-0800."
}
],
"primaryPhone": "+16509030800",
"socials": {
"linkedin": [
{
"url": "https://www.linkedin.com/company/mozilla-corporation/",
"handle": "mozilla-corporation",
"sourceUrl": "https://www.mozilla.org/"
}
],
"instagram": [
{ "url": "https://www.instagram.com/mozilla/", "handle": "mozilla", "sourceUrl": "https://www.mozilla.org/" },
{ "url": "https://www.instagram.com/firefox/", "handle": "firefox", "sourceUrl": "https://www.mozilla.org/" }
],
"tiktok": [
{ "url": "https://www.tiktok.com/@mozilla", "handle": "mozilla", "sourceUrl": "https://www.mozilla.org/" }
]
}
}

That run found 5 emails, 1 phone number and 8 social profiles across 5 networks.

What this Actor does that others do not

An unreadable domain costs you nothing. If not a single page of a domain can be loaded, no row is written for it — and since you are billed per result row, nothing is charged. The reason is written to the run log (1 domain(s) could not be read and were not charged: …) and to the FAILED_DOMAINS record in the run's key-value store. Most contact scrapers return an empty row with an error message in it, and bill you for it.

Emails hidden by Cloudflare are decoded. When a site runs Cloudflare Email Obfuscation, the address is not in the page source at all: contact@example.com becomes a link to /cdn-cgi/l/email-protection and the address is stored encoded. This Actor decodes it and returns it like any other email. Scrapers that only read the HTML return nothing on those pages — and often report the obfuscation link itself as a broken page.

Every value says where it came from. Each email, phone and profile carries the sourceUrl of the page it was found on; emails and phones also carry a snippet of the surrounding text. You can check any value in one click instead of trusting a list.

Phone numbers come out in E.164. +1 650-903-0800 is returned both raw (valueRaw) and normalised (valueE164, here +16509030800), with the country inferred from the site when it is not written in the number. Dates, VAT numbers, French SIRET numbers and GPS coordinates are filtered out.

Social profiles are sorted by network, and matched on the domain name. You get a stable field per network — one CSV column each — not a flat list of links to sort out yourself. The matching is done on the host name, so a link to firefox.com is not reported as an X profile and a ?next=instagram.com/… parameter does not turn an internal link into a social profile.

One row per domain, always. No matter how many pages were read, you get exactly one record per site, so the result is directly usable as a spreadsheet of companies.

Input

FieldTypeDefaultWhat it does
startUrlslist of URLs—The sites to process, one per domain. Also accepts plain text, one URL per line. www.example.com and example.com in the same list are merged into one.
timeoutSecsinteger, 5–12030Timeout for a single HTTP request. Raise it for slow sites, lower it to move faster through a long list.
includeContactsbooleantrueExtract emails and phone numbers. Turn off to collect social profiles only.
includeSocialsbooleantrueExtract social media profiles. Turn off to collect emails and phones only.
{
"startUrls": [
{ "url": "https://www.mozilla.org" },
{ "url": "https://www.debian.org" }
],
"timeoutSecs": 30,
"includeContacts": true,
"includeSocials": true
}

Output fields

Fields that would be empty are left out of the record rather than returned empty, so an export never carries dead columns.

FieldAlways presentContent
domainyesRegistrable domain, e.g. mozilla.org
finalUrlyesURL actually reached, after redirects
keyPagesyesContact, about, legal and privacy pages found on the site
pagesVisitedyesEvery page read for this domain
emailswhen includeContactsOne entry per email: value, type, priority, signals, sourceUrl, snippet, foundIn
primaryEmailwhen one was foundThe single best email for this site, as a plain string
phoneswhen includeContactsOne entry per number: valueRaw, valueE164, priority, signals, sourceUrl, snippet
primaryPhonewhen one was foundThe single best phone number, E.164 when it could be normalised
socialswhen at least one profile was foundProfiles grouped by network: linkedin, facebook, instagram, x, tiktok, youtube, pinterest, google (Google Maps). Each entry has url, handle, sourceUrl
errorsonly when a page failedThe pages that could not be read on an otherwise readable site

type classifies an email as general, sales, support, booking, press or billing. primaryEmail prefers an address on the site's own domain, then a mailto: link, then the contact page. primaryPhone prefers a number from the footer or the contact page, then a tel: link.

How the crawl works

The Actor starts from the URL you give it and looks for the pages where contact details actually live: /contact, /about, /legal, /imprint, /privacy and their French equivalents (/nous-contacter, /a-propos, /mentions-legales, …).

It reads up to 8 pages per site. It goes up to 15 on its own when the site is heavily structured (four or more key pages found) or when the first pages produced neither an email nor a phone number. Assets, PDFs and images are never fetched.

If the URL fails, the Actor retries the obvious variants by itself: http/https, www/no www. Timeouts, network errors, 429 and 5xx responses are retried; a 404 is not.

What you pay for

You are billed one result per domain successfully read, whatever the number of pages that were needed to read it. A domain that could not be read at all produces no result and is not billed. The current rate is shown in the Pricing section of this page.

Limits, stated plainly

  • No proxy. Sites that block datacenter IP ranges — some large marketplaces and review sites do — cannot be read from here. They return no row, so they cost you nothing, and they are listed in FAILED_DOMAINS.
  • Pages are read as HTML, not rendered. Contact details that only appear after JavaScript runs are not seen. Cloudflare-obfuscated emails are the exception: they are decoded.
  • Emails are not verified. No MX or SMTP check is performed; the Actor returns what the site publishes, with the page it was published on.
  • No OCR. An address written inside an image is not read.
  • Domains are processed one after another, in about 1 to 5 seconds each. There is no cap on the number of domains in a run, but the run's own timeout is the real ceiling on a long list — raise it before submitting hundreds of sites.
  • One row per registrable domain. Subdomains are not crawled separately.

Typical uses

  • Enrich a list of companies with a contact email and a phone number before outreach.
  • Build a per-domain social profile list, one column per network, for a partnership or influencer shortlist.
  • Audit what contact details your own sites — or your clients' — expose publicly.