Website Contact Scraper - Email, Phone & Social Media Extractor
Pricing
$10.00 / 1,000 results
Website Contact Scraper - Email, Phone & Social Media Extractor
Extract contact data from any website: emails, phone numbers and social profiles. One row per domain, each email typed and traced to the page it came from, phones in E.164, socials sorted by network. Crawls up to 15 pages per site, starting with contact, about, legal and privacy pages.
Pricing
$10.00 / 1,000 results
Rating
0.0
(0)
Developer
My Smart Digital
Maintained by CommunityActor stats
2
Bookmarked
25
Total users
5
Monthly active users
2 days ago
Last modified
Share
Website Contact Scraper – Email, Phone & Social Media Extractor
Give this Actor a list of website URLs. For each one you get a single row holding the emails, phone numbers and social media profiles published on that site — every value carrying the exact page it was found on.
Feed it the domains you already have and export one CSV where one row is one company: ready for a CRM import, a prospecting list, or an audit of what your own sites expose publicly.
No browser, no proxy, no configuration: a site is read in about 1 to 5 seconds.
One row, for real
Below is an actual result for https://www.mozilla.org, shortened to keep it readable — 15 pages
read in 3.2 seconds. Nothing here is an illustration: these are the fields and the values the Actor
returned.
{"domain": "mozilla.org","finalUrl": "https://www.mozilla.org/en-US/","keyPages": {"contact": "https://www.mozilla.org/en-US/contact/","about": "https://www.mozilla.org/en-US/about/","legal": "https://www.mozilla.org/en-US/about/legal/","privacy": "https://www.mozilla.org/en-US/privacy/websites/cookie-settings/"},"pagesVisited": ["https://www.mozilla.org/en-US/","https://www.mozilla.org/en-US/contact/","https://www.mozilla.org/en-US/about/legal/report-infringement/"],"emails": [{"value": "licensing@mozilla.org","type": "general","priority": "primary","signals": ["text", "same_domain"],"sourceUrl": "https://www.mozilla.org/en-US/about/legal/terms/mozilla","snippet": "tions on Mozilla authored content can be sent to: licensing@mozilla.org.","foundIn": "text"},{"value": "dmcanotice@mozilla.com","type": "general","priority": "secondary","signals": ["mailto"],"sourceUrl": "https://www.mozilla.org/en-US/about/legal/report-infringement","snippet": "If you prefer, you can email dmcanotice@mozilla.com.","foundIn": "mailto"}],"primaryEmail": "licensing@mozilla.org","phones": [{"valueRaw": "+1 650-903-0800","valueE164": "+16509030800","priority": "primary","signals": ["text"],"sourceUrl": "https://www.mozilla.org/en-US/about/legal/report-infringement","snippet": "Mozilla's Designated Agent's phone is +1 650-903-0800."}],"primaryPhone": "+16509030800","socials": {"linkedin": [{"url": "https://www.linkedin.com/company/mozilla-corporation/","handle": "mozilla-corporation","sourceUrl": "https://www.mozilla.org/"}],"instagram": [{ "url": "https://www.instagram.com/mozilla/", "handle": "mozilla", "sourceUrl": "https://www.mozilla.org/" },{ "url": "https://www.instagram.com/firefox/", "handle": "firefox", "sourceUrl": "https://www.mozilla.org/" }],"tiktok": [{ "url": "https://www.tiktok.com/@mozilla", "handle": "mozilla", "sourceUrl": "https://www.mozilla.org/" }]}}
That run found 5 emails, 1 phone number and 8 social profiles across 5 networks.
What this Actor does that others do not
An unreadable domain costs you nothing. If not a single page of a domain can be loaded, no row
is written for it — and since you are billed per result row, nothing is charged. The reason is
written to the run log (1 domain(s) could not be read and were not charged: …) and to the
FAILED_DOMAINS record in the run's key-value store. Most contact scrapers return an empty row
with an error message in it, and bill you for it.
Emails hidden by Cloudflare are decoded. When a site runs Cloudflare Email Obfuscation, the
address is not in the page source at all: contact@example.com becomes a link to
/cdn-cgi/l/email-protection and the address is stored encoded. This Actor decodes it and returns
it like any other email. Scrapers that only read the HTML return nothing on those pages — and often
report the obfuscation link itself as a broken page.
Every value says where it came from. Each email, phone and profile carries the sourceUrl of
the page it was found on; emails and phones also carry a snippet of the surrounding text. You can
check any value in one click instead of trusting a list.
Phone numbers come out in E.164. +1 650-903-0800 is returned both raw (valueRaw) and
normalised (valueE164, here +16509030800), with the country inferred from the site when it is
not written in the number. Dates, VAT numbers, French SIRET numbers and GPS coordinates are
filtered out.
Social profiles are sorted by network, and matched on the domain name. You get a stable field
per network — one CSV column each — not a flat list of links to sort out yourself. The matching is
done on the host name, so a link to firefox.com is not reported as an X profile and a
?next=instagram.com/… parameter does not turn an internal link into a social profile.
One row per domain, always. No matter how many pages were read, you get exactly one record per site, so the result is directly usable as a spreadsheet of companies.
Input
| Field | Type | Default | What it does |
|---|---|---|---|
startUrls | list of URLs | — | The sites to process, one per domain. Also accepts plain text, one URL per line. www.example.com and example.com in the same list are merged into one. |
timeoutSecs | integer, 5–120 | 30 | Timeout for a single HTTP request. Raise it for slow sites, lower it to move faster through a long list. |
includeContacts | boolean | true | Extract emails and phone numbers. Turn off to collect social profiles only. |
includeSocials | boolean | true | Extract social media profiles. Turn off to collect emails and phones only. |
{"startUrls": [{ "url": "https://www.mozilla.org" },{ "url": "https://www.debian.org" }],"timeoutSecs": 30,"includeContacts": true,"includeSocials": true}
Output fields
Fields that would be empty are left out of the record rather than returned empty, so an export never carries dead columns.
| Field | Always present | Content |
|---|---|---|
domain | yes | Registrable domain, e.g. mozilla.org |
finalUrl | yes | URL actually reached, after redirects |
keyPages | yes | Contact, about, legal and privacy pages found on the site |
pagesVisited | yes | Every page read for this domain |
emails | when includeContacts | One entry per email: value, type, priority, signals, sourceUrl, snippet, foundIn |
primaryEmail | when one was found | The single best email for this site, as a plain string |
phones | when includeContacts | One entry per number: valueRaw, valueE164, priority, signals, sourceUrl, snippet |
primaryPhone | when one was found | The single best phone number, E.164 when it could be normalised |
socials | when at least one profile was found | Profiles grouped by network: linkedin, facebook, instagram, x, tiktok, youtube, pinterest, google (Google Maps). Each entry has url, handle, sourceUrl |
errors | only when a page failed | The pages that could not be read on an otherwise readable site |
type classifies an email as general, sales, support, booking, press or billing.
primaryEmail prefers an address on the site's own domain, then a mailto: link, then the contact
page. primaryPhone prefers a number from the footer or the contact page, then a tel: link.
How the crawl works
The Actor starts from the URL you give it and looks for the pages where contact details actually
live: /contact, /about, /legal, /imprint, /privacy and their French equivalents
(/nous-contacter, /a-propos, /mentions-legales, …).
It reads up to 8 pages per site. It goes up to 15 on its own when the site is heavily structured (four or more key pages found) or when the first pages produced neither an email nor a phone number. Assets, PDFs and images are never fetched.
If the URL fails, the Actor retries the obvious variants by itself: http/https, www/no www.
Timeouts, network errors, 429 and 5xx responses are retried; a 404 is not.
What you pay for
You are billed one result per domain successfully read, whatever the number of pages that were needed to read it. A domain that could not be read at all produces no result and is not billed. The current rate is shown in the Pricing section of this page.
Limits, stated plainly
- No proxy. Sites that block datacenter IP ranges — some large marketplaces and review sites do
— cannot be read from here. They return no row, so they cost you nothing, and they are listed in
FAILED_DOMAINS. - Pages are read as HTML, not rendered. Contact details that only appear after JavaScript runs are not seen. Cloudflare-obfuscated emails are the exception: they are decoded.
- Emails are not verified. No MX or SMTP check is performed; the Actor returns what the site publishes, with the page it was published on.
- No OCR. An address written inside an image is not read.
- Domains are processed one after another, in about 1 to 5 seconds each. There is no cap on the number of domains in a run, but the run's own timeout is the real ceiling on a long list — raise it before submitting hundreds of sites.
- One row per registrable domain. Subdomains are not crawled separately.
Typical uses
- Enrich a list of companies with a contact email and a phone number before outreach.
- Build a per-domain social profile list, one column per network, for a partnership or influencer shortlist.
- Audit what contact details your own sites — or your clients' — expose publicly.