Website Email Scraper – Contact Details, Socials & WhatsApp
Pricing
from $1.20 / 1,000 results
Website Email Scraper – Contact Details, Socials & WhatsApp
Contact details scraper for any list of website URLs: role emails, social profiles, WhatsApp click-to-chat numbers, the contact page and tech stack, plus the page as clean Markdown or text from the same fetch. No browser, robots.txt honoured.
Pricing
from $1.20 / 1,000 results
Rating
0.0
(0)
Developer
Locomint
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
3
Monthly active users
5 days ago
Last modified
Categories
Share
This contact details scraper reads a list of website URLs and returns, for each page, the email addresses, social profiles, WhatsApp number, contact page and tech stack it publishes, together with the page itself as clean Markdown or plain text. Both come from one HTTP fetch and cost one price per page.
What it does
For every URL you give it, the actor checks the site's robots.txt, fetches the page over plain HTTP (no browser) and writes one dataset row. The row carries two things.
Contact points found in the page's HTML
| Field | What goes in it |
|---|---|
emails | Every address the page publishes: up to five company mailboxes such as info@, sales@ and bookings@ first, then up to five addresses that name a member of staff. Addresses anywhere in the HTML are read, including mailto: links and entity-encoded ones. |
named_emails | The addresses in emails that name a person (maria.silva@, jsmith@), on the site's own domain or on a free mail provider. Remove them from emails to keep company mailboxes only. |
socials | The first profile link per network: Facebook, Instagram, LinkedIn, X (Twitter), YouTube and TikTok. Share buttons, intent links and single posts are skipped, and linkedin is always the company's own page. |
named_profiles | Up to five personal profile links the page publishes for its staff, such as linkedin.com/in/maria-silva. Only the link is returned; the profile is never opened. |
whatsapp | The number in the first WhatsApp click-to-chat link (wa.me, api.whatsapp.com/send, whatsapp://send), written as + and digits. |
contact_form_url | The first link on the same site whose path contains "contact". |
tech_stack | Technologies matched by 36 page signatures (WordPress, WooCommerce, Shopify, Wix, Squarespace, Webflow, HubSpot, Intercom, Google Tag Manager, Stripe and others) plus four read from response headers (Cloudflare, PHP, ASP.NET, Express). |
The page as text
content holds the page body as Markdown, with headings, lists, links, quotes and code blocks
kept, or as plain text, cut at maxChars. Scripts, styles, forms, iframes and the nav,
header, footer and aside blocks are removed first. Next to it are title,
description (the meta or Open Graph description), language (as declared in
<html lang>), canonical_url, up to 50 headings, word_count and truncated.
Every URL comes back as a row with a status, so a page that failed is visible in the dataset
instead of missing from it. Contacts and text are one page event between them; turning
contacts off does not make a page cheaper.
Who it is for
- Lead generation and sales prospecting. Run the website column of a prospect list, or the websites returned by the Locomint Google Maps actors, and get each company's published mailboxes, WhatsApp number and social profiles back as columns.
- Agencies building contact sheets for a niche or a city. The tech stack sits in the same row, so "Shopify stores with no WhatsApp link" is a spreadsheet filter rather than an afternoon of clicking through sites.
- CRM enrichment. Export the website field, run it, and merge
emails,socialsandcontact_form_urlback onurl.final_urlshows where each site redirected, which catches companies that have moved domain. - RAG and content teams who need a known list of pages as clean Markdown with metadata. The contacts arrive in the same row at no extra cost.
How to use it
In the Apify Console:
- Paste full URLs, including
https://, into Page URLs, one per line, up to 5,000. The input form rejects bare domains such asexample.com. - Leave Extract business contacts on and choose Markdown or Plain text. If you only want the contacts, set Max characters per page to 500 to keep the dataset small; the whole page is still read for contacts.
- Set a maximum cost per run if you want a ceiling. The actor stops before the row that would pass it.
- Start the run. The Pages view shows URL, status, title, word count, language, emails, WhatsApp and tech stack, and the Content view shows the text. Export as CSV, Excel or JSON.
From the API, this call starts a run, waits for it and returns the rows:
curl -X POST \"https://api.apify.com/v2/acts/locomint~website-content-contact-extractor/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls": [{"url": "https://www.python.org/"}], "includeContacts": true}'
The synchronous endpoint waits up to 300 seconds. For a long list, start the run with
POST https://api.apify.com/v2/acts/locomint~website-content-contact-extractor/runs?token=YOUR_APIFY_TOKEN
and read its dataset when the run finishes. The official apify-client packages for Python
and JavaScript do the same in a few lines.
Input example
{"startUrls": [{ "url": "https://www.python.org/" },{ "url": "https://docs.python.org/3/" }],"includeContacts": true,"outputFormat": "markdown","maxChars": 20000,"includeLinks": false,"concurrency": 5}
| Field | Default | Allowed | What it does |
|---|---|---|---|
startUrls | required | 1 to 5,000 URLs | The pages to read. Exact duplicates are dropped; over 5,000, the first 5,000 are used. |
includeContacts | true | true / false | Emails, socials, WhatsApp, contact page and tech stack from the same fetch. |
outputFormat | markdown | markdown, text | Markdown keeps structure and links; text is the words only. |
maxChars | 50,000 | 500 to 500,000 | Longer content is cut, ends with [truncated] and sets truncated: true. |
includeLinks | false | true / false | Adds links: up to 500 absolute URLs found on the page. |
concurrency | 5 | 1 to 10 | Pages fetched at the same time. |
Through the API, startUrls also accepts plain strings instead of { "url": ... } objects.
Output example
A real row, from a run on https://www.python.org/ on 11 September 2026. content and
headings are trimmed here; the row held 4,624 characters of content and 9 headings.
python.org publishes no mailbox on its homepage, so emails is empty; on a business site
that lists info@ or sales@, the addresses appear in that list.
{"url": "https://www.python.org/","final_url": "https://www.python.org/","status": "ok","http_status": 200,"title": "Welcome to Python.org","description": "The official home of the Python Programming Language","language": "en","canonical_url": null,"headings": ["Get Started", "Download", "Docs", "Jobs", "Latest News", "Upcoming Events"],"content": "[...]\n\n## Get Started\n\nWhether you're new to programming or an experienced developer, [...]\n\n[Start with our Beginner’s Guide](https://www.python.org/about/gettingstarted/)\n\n## Download\n\n[...]","word_count": 332,"truncated": false,"fetched_at": "2026-09-11T18:16:05+00:00","emails": [],"socials": {"linkedin": "https://www.linkedin.com/company/python-software-foundation","x": "https://twitter.com/ThePSF"},"whatsapp": null,"contact_form_url": null,"tech_stack": ["jquery", "font-awesome"]}
Two fields appear only when they have something to say: links (with includeLinks on), and
note, which explains an unusual fetch: robots.txt disallowed the path (with the rule),
robots.txt could not be read so no rules were applied, or the page gave no answer within 25
seconds.
Status values
status | Meaning | Content and contacts |
|---|---|---|
ok | Fetched and read. | Filled. |
thin | Fewer than 40 words of server-rendered text, which usually means the site builds its pages with JavaScript. | Whatever the HTML holds; contact links in the raw HTML are still found. |
parked | The page matches a parked-domain, for-sale, coming-soon or under-construction signature. | Filled from what the page says. |
http_error | The site answered 4xx or 5xx; http_status has the code. 5xx answers are tried three times first. | Empty. |
not_html | The URL is a PDF, image or other non-text file. | Empty. |
unreachable | No connection, a DNS failure, or no answer within 25 seconds. | Empty. |
blocked | The site answered with 403, 429 or a challenge page. Not retried. | Empty. |
refused | The URL points at a private or internal network address. | Empty. |
robots_disallowed | The site's robots.txt disallows the path, so the page was not fetched. | Empty. |
Pricing
| Event | Price |
|---|---|
| Page delivered (one dataset row) | $0.0008, which is $0.80 per 1,000 until 26 September 2026, then $0.0015, which is $1.50 per 1,000 |
| Actor start | $0.00005 per GB of run memory, charged once per run |
Worked example: the homepages of 5,000 companies cost 5,000 x $0.0008 = $4.00, plus $0.00005 for starting a 1 GB run. Reading both the homepage and the contact page of each company is 10,000 pages. From 26 September 2026 the price per page becomes $0.0015.
You pay only these event prices; Apify compute is not billed to you separately, and contacts
cost nothing on top of the page. Every URL in the list produces one row and one charge,
including rows whose status is unreachable, blocked or robots_disallowed, because each is
an answer about that URL. If you set a maximum cost per run, the actor stops before the row
that would pass it.
FAQ
Does it render JavaScript?
No. It reads the HTML the server sends, which is why it is fast and cheap per page.
Sites that build their content in the browser come back thin; their contact links are often
still in the raw HTML and are still extracted.
Why is an email missing that I can see on the site?
The address may be on another page, since this actor reads only the URLs you send. It may be
inserted by JavaScript, hidden behind Cloudflare's email protection, written as
"info [at] example.com" or shown as an image, none of which is decoded. Or it names a person
on another company's domain, such as the web designer credited in the footer, which is not
returned. An unnamed address on another domain is kept only when it names a business
function, noreply@ and postmaster@ addresses are skipped, and at most five company
mailboxes and five named addresses are returned per page.
Does it follow links to the contact page?
No. It returns exactly one row per URL you send, which keeps the cost predictable. Run the
homepages first, then send the contact_form_url values as a second list. To walk a whole
site, use the Website Content Crawler listed below.
Does it extract phone numbers?
Only WhatsApp numbers, from click-to-chat links, returned as + and digits. Other phone
numbers written on the page are not parsed. The Locomint Google Maps actors return each
business's listed phone number.
Does it honour robots.txt?
Yes. It reads each site's /robots.txt once per run and applies the rules for all crawlers
(the * group). A disallowed URL is not fetched and comes back as robots_disallowed, with
the matching rule in note. If robots.txt cannot be read, the page is fetched and note says
no rules were applied; a site with no robots.txt has no rules.
What happens when a site blocks it?
A 403, a 429 or a challenge page is reported as blocked and never retried, from the same
address or any other. The site is then left alone for 15 minutes: other URLs on the same host
in that window come back blocked without a request being made. The actor does not solve
CAPTCHAs or switch IP addresses to get past a block.
How is the tech stack detected?
By matching 36 signatures against the page's HTML, such as /wp-content/ for WordPress or
cdn.shopify.com for Shopify, and by reading response headers for Cloudflare, PHP, ASP.NET
and Express. It reports what the page loads, so a script a theme includes but never uses
still counts.
Limits
- No JavaScript rendering. Single-page apps come back
thinand their text is missing. - One page per URL, 5,000 URLs per run, 25 seconds per URL, 3 MB of HTML per page, and at most 5 redirects, each checked before it is followed.
- Up to five company mailboxes and five named addresses per page. A named address on another company's domain is not returned. Cloudflare-protected, image and "[at]" addresses are not decoded.
- Phone numbers: WhatsApp click-to-chat links only.
- Social profiles: the first link per network on the page, plus up to five personal profile
links in
named_profiles. The actor records the link and never opens the profile. - Boilerplate is removed by tag. Navigation built from plain
divelements, as on documentation sites made with Sphinx, stays incontent. languageis what the page declares, not a detection from the text, andnullwhen the page declares nothing.- No raw HTML in the output.
Compliance
Contact data comes only from what the site publishes. That includes addresses that name a
member of staff and links to staff profiles, which are kept in named_emails and
named_profiles so you can use them or leave them out; a profile page is never opened. A
named work address is personal data in the EU and the UK. You are responsible for using the
data lawfully, including whether you may write to a named person (for example under GDPR,
CAN-SPAM, PECR and local law).
Questions, bug reports and feature requests go on this actor's Issues tab. Anyone can ask for their own address or profile link to be left out, and a business can ask for its whole site to be left out, by writing to info@locomint.io; that address is for removal requests only, and removals are honoured in every Locomint listing. This actor keeps no copy of what it reads between runs, so each run returns what the public pages show at that moment.
Other Locomint actors
- Google Maps Scraper & Email Extractor – Business Leads: Search terms and a city in, business records with website contacts out.
- Google Maps Scraper – Multi-City Lead Lists with Emails: Many categories across many cities in one deduplicated run.
- Google Maps Place Details Scraper – Bulk Place ID Lookup: Place IDs or place-page links in, full records out.
- Website Content Crawler – Markdown for AI, Emails & Contacts: A whole site as Markdown, with its contact points.
- Bulk Email Verifier & Validator: Checks whether addresses can receive mail.
- Company Enrichment API – Domain to Emails, Socials & Tech: A domain in, its contacts and technologies out.
- AI Crawler Checker – robots.txt Rules for GPTBot & ClaudeBot: Which AI crawlers a site's robots.txt allows.
- Schema Markup Validator & Generator – JSON-LD Checker: Checks and generates schema.org markup.