Website Email Scraper – Contact Details, Socials & WhatsApp avatar

Website Email Scraper – Contact Details, Socials & WhatsApp

Pricing

from $1.20 / 1,000 results

Go to Apify Store
Website Email Scraper – Contact Details, Socials & WhatsApp

Website Email Scraper – Contact Details, Socials & WhatsApp

Contact details scraper for any list of website URLs: role emails, social profiles, WhatsApp click-to-chat numbers, the contact page and tech stack, plus the page as clean Markdown or text from the same fetch. No browser, robots.txt honoured.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Locomint

Locomint

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

3

Monthly active users

5 days ago

Last modified

Share

This contact details scraper reads a list of website URLs and returns, for each page, the email addresses, social profiles, WhatsApp number, contact page and tech stack it publishes, together with the page itself as clean Markdown or plain text. Both come from one HTTP fetch and cost one price per page.

What it does

For every URL you give it, the actor checks the site's robots.txt, fetches the page over plain HTTP (no browser) and writes one dataset row. The row carries two things.

Contact points found in the page's HTML

FieldWhat goes in it
emailsEvery address the page publishes: up to five company mailboxes such as info@, sales@ and bookings@ first, then up to five addresses that name a member of staff. Addresses anywhere in the HTML are read, including mailto: links and entity-encoded ones.
named_emailsThe addresses in emails that name a person (maria.silva@, jsmith@), on the site's own domain or on a free mail provider. Remove them from emails to keep company mailboxes only.
socialsThe first profile link per network: Facebook, Instagram, LinkedIn, X (Twitter), YouTube and TikTok. Share buttons, intent links and single posts are skipped, and linkedin is always the company's own page.
named_profilesUp to five personal profile links the page publishes for its staff, such as linkedin.com/in/maria-silva. Only the link is returned; the profile is never opened.
whatsappThe number in the first WhatsApp click-to-chat link (wa.me, api.whatsapp.com/send, whatsapp://send), written as + and digits.
contact_form_urlThe first link on the same site whose path contains "contact".
tech_stackTechnologies matched by 36 page signatures (WordPress, WooCommerce, Shopify, Wix, Squarespace, Webflow, HubSpot, Intercom, Google Tag Manager, Stripe and others) plus four read from response headers (Cloudflare, PHP, ASP.NET, Express).

The page as text

content holds the page body as Markdown, with headings, lists, links, quotes and code blocks kept, or as plain text, cut at maxChars. Scripts, styles, forms, iframes and the nav, header, footer and aside blocks are removed first. Next to it are title, description (the meta or Open Graph description), language (as declared in <html lang>), canonical_url, up to 50 headings, word_count and truncated.

Every URL comes back as a row with a status, so a page that failed is visible in the dataset instead of missing from it. Contacts and text are one page event between them; turning contacts off does not make a page cheaper.

Who it is for

  • Lead generation and sales prospecting. Run the website column of a prospect list, or the websites returned by the Locomint Google Maps actors, and get each company's published mailboxes, WhatsApp number and social profiles back as columns.
  • Agencies building contact sheets for a niche or a city. The tech stack sits in the same row, so "Shopify stores with no WhatsApp link" is a spreadsheet filter rather than an afternoon of clicking through sites.
  • CRM enrichment. Export the website field, run it, and merge emails, socials and contact_form_url back on url. final_url shows where each site redirected, which catches companies that have moved domain.
  • RAG and content teams who need a known list of pages as clean Markdown with metadata. The contacts arrive in the same row at no extra cost.

How to use it

In the Apify Console:

  1. Paste full URLs, including https://, into Page URLs, one per line, up to 5,000. The input form rejects bare domains such as example.com.
  2. Leave Extract business contacts on and choose Markdown or Plain text. If you only want the contacts, set Max characters per page to 500 to keep the dataset small; the whole page is still read for contacts.
  3. Set a maximum cost per run if you want a ceiling. The actor stops before the row that would pass it.
  4. Start the run. The Pages view shows URL, status, title, word count, language, emails, WhatsApp and tech stack, and the Content view shows the text. Export as CSV, Excel or JSON.

From the API, this call starts a run, waits for it and returns the rows:

curl -X POST \
"https://api.apify.com/v2/acts/locomint~website-content-contact-extractor/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls": [{"url": "https://www.python.org/"}], "includeContacts": true}'

The synchronous endpoint waits up to 300 seconds. For a long list, start the run with POST https://api.apify.com/v2/acts/locomint~website-content-contact-extractor/runs?token=YOUR_APIFY_TOKEN and read its dataset when the run finishes. The official apify-client packages for Python and JavaScript do the same in a few lines.

Input example

{
"startUrls": [
{ "url": "https://www.python.org/" },
{ "url": "https://docs.python.org/3/" }
],
"includeContacts": true,
"outputFormat": "markdown",
"maxChars": 20000,
"includeLinks": false,
"concurrency": 5
}
FieldDefaultAllowedWhat it does
startUrlsrequired1 to 5,000 URLsThe pages to read. Exact duplicates are dropped; over 5,000, the first 5,000 are used.
includeContactstruetrue / falseEmails, socials, WhatsApp, contact page and tech stack from the same fetch.
outputFormatmarkdownmarkdown, textMarkdown keeps structure and links; text is the words only.
maxChars50,000500 to 500,000Longer content is cut, ends with [truncated] and sets truncated: true.
includeLinksfalsetrue / falseAdds links: up to 500 absolute URLs found on the page.
concurrency51 to 10Pages fetched at the same time.

Through the API, startUrls also accepts plain strings instead of { "url": ... } objects.

Output example

A real row, from a run on https://www.python.org/ on 11 September 2026. content and headings are trimmed here; the row held 4,624 characters of content and 9 headings. python.org publishes no mailbox on its homepage, so emails is empty; on a business site that lists info@ or sales@, the addresses appear in that list.

{
"url": "https://www.python.org/",
"final_url": "https://www.python.org/",
"status": "ok",
"http_status": 200,
"title": "Welcome to Python.org",
"description": "The official home of the Python Programming Language",
"language": "en",
"canonical_url": null,
"headings": ["Get Started", "Download", "Docs", "Jobs", "Latest News", "Upcoming Events"],
"content": "[...]\n\n## Get Started\n\nWhether you're new to programming or an experienced developer, [...]\n\n[Start with our Beginner’s Guide](https://www.python.org/about/gettingstarted/)\n\n## Download\n\n[...]",
"word_count": 332,
"truncated": false,
"fetched_at": "2026-09-11T18:16:05+00:00",
"emails": [],
"socials": {
"linkedin": "https://www.linkedin.com/company/python-software-foundation",
"x": "https://twitter.com/ThePSF"
},
"whatsapp": null,
"contact_form_url": null,
"tech_stack": ["jquery", "font-awesome"]
}

Two fields appear only when they have something to say: links (with includeLinks on), and note, which explains an unusual fetch: robots.txt disallowed the path (with the rule), robots.txt could not be read so no rules were applied, or the page gave no answer within 25 seconds.

Status values

statusMeaningContent and contacts
okFetched and read.Filled.
thinFewer than 40 words of server-rendered text, which usually means the site builds its pages with JavaScript.Whatever the HTML holds; contact links in the raw HTML are still found.
parkedThe page matches a parked-domain, for-sale, coming-soon or under-construction signature.Filled from what the page says.
http_errorThe site answered 4xx or 5xx; http_status has the code. 5xx answers are tried three times first.Empty.
not_htmlThe URL is a PDF, image or other non-text file.Empty.
unreachableNo connection, a DNS failure, or no answer within 25 seconds.Empty.
blockedThe site answered with 403, 429 or a challenge page. Not retried.Empty.
refusedThe URL points at a private or internal network address.Empty.
robots_disallowedThe site's robots.txt disallows the path, so the page was not fetched.Empty.

Pricing

EventPrice
Page delivered (one dataset row)$0.0008, which is $0.80 per 1,000 until 26 September 2026, then $0.0015, which is $1.50 per 1,000
Actor start$0.00005 per GB of run memory, charged once per run

Worked example: the homepages of 5,000 companies cost 5,000 x $0.0008 = $4.00, plus $0.00005 for starting a 1 GB run. Reading both the homepage and the contact page of each company is 10,000 pages. From 26 September 2026 the price per page becomes $0.0015.

You pay only these event prices; Apify compute is not billed to you separately, and contacts cost nothing on top of the page. Every URL in the list produces one row and one charge, including rows whose status is unreachable, blocked or robots_disallowed, because each is an answer about that URL. If you set a maximum cost per run, the actor stops before the row that would pass it.

FAQ

Does it render JavaScript?

No. It reads the HTML the server sends, which is why it is fast and cheap per page. Sites that build their content in the browser come back thin; their contact links are often still in the raw HTML and are still extracted.

Why is an email missing that I can see on the site?

The address may be on another page, since this actor reads only the URLs you send. It may be inserted by JavaScript, hidden behind Cloudflare's email protection, written as "info [at] example.com" or shown as an image, none of which is decoded. Or it names a person on another company's domain, such as the web designer credited in the footer, which is not returned. An unnamed address on another domain is kept only when it names a business function, noreply@ and postmaster@ addresses are skipped, and at most five company mailboxes and five named addresses are returned per page.

No. It returns exactly one row per URL you send, which keeps the cost predictable. Run the homepages first, then send the contact_form_url values as a second list. To walk a whole site, use the Website Content Crawler listed below.

Does it extract phone numbers?

Only WhatsApp numbers, from click-to-chat links, returned as + and digits. Other phone numbers written on the page are not parsed. The Locomint Google Maps actors return each business's listed phone number.

Does it honour robots.txt?

Yes. It reads each site's /robots.txt once per run and applies the rules for all crawlers (the * group). A disallowed URL is not fetched and comes back as robots_disallowed, with the matching rule in note. If robots.txt cannot be read, the page is fetched and note says no rules were applied; a site with no robots.txt has no rules.

What happens when a site blocks it?

A 403, a 429 or a challenge page is reported as blocked and never retried, from the same address or any other. The site is then left alone for 15 minutes: other URLs on the same host in that window come back blocked without a request being made. The actor does not solve CAPTCHAs or switch IP addresses to get past a block.

How is the tech stack detected?

By matching 36 signatures against the page's HTML, such as /wp-content/ for WordPress or cdn.shopify.com for Shopify, and by reading response headers for Cloudflare, PHP, ASP.NET and Express. It reports what the page loads, so a script a theme includes but never uses still counts.

Limits

  • No JavaScript rendering. Single-page apps come back thin and their text is missing.
  • One page per URL, 5,000 URLs per run, 25 seconds per URL, 3 MB of HTML per page, and at most 5 redirects, each checked before it is followed.
  • Up to five company mailboxes and five named addresses per page. A named address on another company's domain is not returned. Cloudflare-protected, image and "[at]" addresses are not decoded.
  • Phone numbers: WhatsApp click-to-chat links only.
  • Social profiles: the first link per network on the page, plus up to five personal profile links in named_profiles. The actor records the link and never opens the profile.
  • Boilerplate is removed by tag. Navigation built from plain div elements, as on documentation sites made with Sphinx, stays in content.
  • language is what the page declares, not a detection from the text, and null when the page declares nothing.
  • No raw HTML in the output.

Compliance

Contact data comes only from what the site publishes. That includes addresses that name a member of staff and links to staff profiles, which are kept in named_emails and named_profiles so you can use them or leave them out; a profile page is never opened. A named work address is personal data in the EU and the UK. You are responsible for using the data lawfully, including whether you may write to a named person (for example under GDPR, CAN-SPAM, PECR and local law).

Questions, bug reports and feature requests go on this actor's Issues tab. Anyone can ask for their own address or profile link to be left out, and a business can ask for its whole site to be left out, by writing to info@locomint.io; that address is for removal requests only, and removals are honoured in every Locomint listing. This actor keeps no copy of what it reads between runs, so each run returns what the public pages show at that moment.

Other Locomint actors