Email Extractor — Website Email, Phone & Social Scraper
Pricing
from $3.50 / 1,000 results
Email Extractor — Website Email, Phone & Social Scraper
Bulk email & phone extractor for a list of websites. Paste URLs or domains, get back deduped emails, phone numbers, and social links per site — homepage plus contact/about pages, junk filtered. For lead lists, CRM enrichment, and outreach.
Pricing
from $3.50 / 1,000 results
Rating
2.0
(1)
Developer
Aitor Sanchez-Mansilla
Maintained by CommunityActor stats
2
Bookmarked
250
Total users
35
Monthly active users
11 days ago
Last modified
Categories
Share
Email Extractor — Website Emails, Phones & Social Links
Extract emails from a list of websites in one run. Give this email extractor a few hundred company or business domains and get back clean, deduped email addresses, phone numbers, and social-media links — ready for lead-gen lists, CRM enrichment, or research.
No setup, no per-site configuration. Give it URLs, get contacts.
Why this email extractor
- Bulk + parallel — feed it hundreds of domains; it fetches many in parallel, so a big batch finishes in minutes, not hours.
- Emails, phones, and socials — not just an address: phone numbers and social links (Instagram, Facebook, LinkedIn, X, YouTube, TikTok) in the same record.
- Finds the contact page instead of guessing it — it reads the site's own navigation and follows the links that lead to contact details, whatever they are called:
/impressum,/kontakt,/mentions-legales,/aviso-legal,/contattior a bespoke path. Conventional paths are only tried as a fallback. - Reads European sites properly — in Germany, Austria and Switzerland the real contact details sit on the legally required imprint page, not on
/contact. Measured on 40 live German sites, following those pages raised the share returning an email from 55% to 82%. - Decodes hidden addresses — many sites encode their email so that the page text only shows
[email protected]. This one decodes it back to the real address. - Best contact first — addresses come back ranked, so
emails[0]is the one worth writing to: the site's own domain and a real inbox (info@,kontakt@,booking@) ahead ofnoreply@andprivacy@. - Reads structured data too — where a site publishes schema.org/JSON-LD, its own declared email, phone and social profiles are picked up as data rather than guessed from text.
- A social profile or Linktree is a valid input — point it at
instagram.com/yourtargetor a link hub and you still get the full set of profiles and any contact details behind them. - Optional Facebook fallback — when a site publishes no email but links a Facebook page, the email is looked up there (measured: works for ~75% of such sites). Enriched rows are marked
recoveredBy: "facebook-page". - Never silently drops a URL — every input returns a row. If a site can't be read you get the row with an
error, so a list of 500 always comes back as 500 records. - Messaging channels too — WhatsApp numbers and group invites, Telegram, Discord, plus Pinterest, Threads and Reddit alongside the usual six networks. For a small operator a WhatsApp line is often the one that gets answered.
- Follows a Linktree to the real sites — point it at a link-in-bio page and it visits the destinations behind it, so you get the email on the actual website, not just the profile links on the hub.
- Cleans up broken addresses — an address welded to the next word by sloppy markup (
…@site.comphone) is repaired rather than returned unusable. - Clean output — deduped and junk-filtered (drops
noreply@, asset filenames, placeholder domains); one tidy record per site. - No per-contact metering, no contract — pay per site scanned, point it at your own list, and keep everything you find.
Email extractor
For each URL you provide, this email extractor:
- Reads the page's published contact details.
- Optionally also checks the same site's /contact, /contact-us, /about, and /about-us pages — where businesses usually publish their email and phone.
- Returns a single tidy record per input URL with every email, phone, and social link it found, deduped.
It processes many pages in parallel, so a big batch of websites finishes in a fraction of the time a one-at-a-time scan would take.
Use cases
- Email extraction at scale — turn a list of website domains into a clean list of email addresses.
- Lead generation — turn a list of business domains into a contact list.
- CRM enrichment — fill in missing email / phone / social fields for accounts you already have.
- Market & competitor research — collect public contact and social presence across a set of sites.
- Outreach prep — find the right email and social handles before reaching out.
Input
| Field | Type | Default | Description |
|---|---|---|---|
urls | array of strings | — (required) | Pages or domains to scan. Bare domains like example.com get https:// added automatically. |
crawlContactPages | boolean | true | Also check each site's contact, about and legal pages (/contact, /about, /impressum…). Best coverage; turn off for a single-page check. |
facebookFallback | boolean | false | If a site publishes no email but links a Facebook page, look the email up there (billed separately, ~$0.01 per page checked). Recovers an email for ~75% of such sites. |
maxItems | integer | unlimited | Cap the number of input URLs processed. |
Example input
{"urls": ["https://acme-coffee.com","blue-fox-studio.com","https://example-agency.com/contact"],"crawlContactPages": true}
Output
One record per input URL:
{"url": "https://acme-coffee.com","finalUrl": "https://acme-coffee.com/","emails": ["hello@acme-coffee.com"],"phones": ["+15551234567"],"socials": {"instagram": ["https://instagram.com/acmecoffee"],"facebook": ["https://facebook.com/acmecoffee"],"twitter": [],"linkedin": ["https://www.linkedin.com/company/acme-coffee"],"youtube": [],"tiktok": []},"otherUrls": ["https://acme-coffee.com/menu"],"pagesScanned": ["https://acme-coffee.com/", "https://acme-coffee.com/contact", "https://acme-coffee.com/about"],"noEmailReason": null,"contactFormUrl": "https://acme-coffee.com/contact","scrapedAt": "2026-06-20T10:00:00.000Z"}
| Field | Description |
|---|---|
url | The URL you supplied. |
finalUrl | Where it landed after redirects. |
emails | Deduped, junk-filtered email addresses (drops noreply@, asset filenames, placeholder domains, etc.). |
phones | Deduped, loosely normalized phone numbers. |
socials | Links grouped by platform: instagram, facebook, twitter, linkedin, youtube, tiktok. |
otherUrls | Other outbound links found on the page (non-social, non-asset). |
pagesScanned | Which pages were actually read for this record. |
noEmailReason | null when an email was found. Otherwise why not: no_email_published, contact_form_only, javascript_only, blocked or unreachable (see Why no email?). |
contactFormUrl | First scanned page with a contact form, or null. Most useful on rows without an email. |
error | Only on rows whose site could not be read: what went wrong. |
scrapedAt | ISO timestamp of the check. |
Cost
Pay-per-result: $4 per 1,000 websites scanned (≈ $0.004 per input URL), with automatic volume discounts down to $2 per 1,000 at higher usage tiers. Platform usage costs are included — the price you see is all you pay — and you're charged once per input URL, regardless of how many contact/about pages it reads for that site.
For comparison: per-contact data tools (Apollo, Hunter, Lusha) charge $0.40–0.80 per verified contact — roughly $400–800 to cover 1,000 sites — and their coverage thins out on smaller, independent domains. This extractor reads what each site already publishes, so 1,000 sites costs a few dollars, not a few hundred. It's a different tool for a different job: bulk public-contact extraction over your list, not metered per-contact lookups.
Real-time HTTP API
Need one site's contacts right now, from an app or an AI agent, without starting a run? The Actor also runs in Standby mode: an always-on HTTP endpoint that answers in a few seconds.
curl -H "Authorization: Bearer $APIFY_TOKEN" \"https://aitorsm--email-extractor.apify.actor/contacts?url=apify.com/contact"
Copy the exact base URL from the Actor's API → Standby tab. Query parameters: url (required), crawlContactPages (default true), maxContactPages (0–8, default 4). The response is the same row as in the output above; a site that cannot be fetched still returns 200 with an error field. Private, loopback and link-local addresses are rejected with 400, and a request that takes longer than 90 seconds returns 504.
AI agents: Use Apify’s MCP connector to call this Actor as a batch run, then fetch its dataset items for the results. For immediate single-site requests, use the Standby HTTP endpoint above.
Cost per call: one result event per returned row (rows with an error included, the same as in a batch run; requests rejected with a 400 are free), plus the compute of the Standby run under your account, idle time included.
facebookFallback is only available in batch runs: it is billed separately and can take many minutes. For lists, use a normal run.
FAQ
How do I extract emails from a list of websites?
Paste your website or domain list into the urls field and run the actor. Each URL is read for published email addresses, phone numbers, and social links, and you get one deduped record per input URL. With crawlContactPages on (the default), it also checks each site's contact and about pages, where contact details are most often published.
Why did a site return no email?
Every row without an email says why in noEmailReason:
| Value | What it means | What to do |
|---|---|---|
no_email_published | The pages loaded and none of them shows an email address. | Nothing more to find on the site; try its socials, or turn on facebookFallback. |
contact_form_only | No address, but the site has a contact form. | Use the page in contactFormUrl. |
javascript_only | The site is an app that draws its content in the browser, so a plain page read sees an empty shell. | Check the site by hand in a browser. |
blocked | The site refused the request: an HTTP 401, 403, 429 or 503, or a bot-challenge page such as "Just a moment…". Also used when the homepage loaded but most of its contact pages were refused. | Try again later, or check it by hand. |
unreachable | The site could not be read at all: the domain does not resolve, the connection timed out or was refused, the TLS handshake failed, or the page answered with an HTTP error. | Check the domain is still live. |
On a sample of 100 event-organizer websites, 62 returned an email; the rest split into the reasons above. contactFormUrl is filled in on any row where a real contact form was found, with or without an email. Newsletter sign-ups, search boxes and login forms are not counted.
Does it work on German, French, Spanish or Italian websites?
Yes. Sites in those markets keep their contact details on an imprint or legal page (/impressum, /mentions-legales, /aviso-legal), not on /contact. German commercial sites are legally required to publish one. On a 40-site German sample, reading those pages took the share of sites returning an email from 55% to 82%.
Why does a site show "[email protected]" instead of an address?
That is email obfuscation: the page ships the address encoded and a script rebuilds it when the page is displayed, so the raw page text only holds the placeholder. This Actor decodes it and returns the real address.
Which email should I actually use?
The first one. Results are ranked: an address on the site's own domain with a real inbox name (info@, kontakt@, office@, booking@) sorts above generic providers and far above noreply@, privacy@ or webmaster@.
Does it get phone numbers and social links too?
Yes. Alongside emails, every record includes deduped phone numbers and social-media links grouped by platform (instagram, facebook, twitter, linkedin, youtube, tiktok), plus any other outbound links found. It's a full website email and contact extractor, not just emails.
What kind of websites work best?
Conventional content and business sites where contact details are published in the page. Pages that only reveal contacts after heavy in-browser loading may show less. All extracted data is public information published on the pages you point it at.
Related Actors
- Eventbrite Scraper — collect organizer websites from events in any city, then extract their emails and phones here.
- Lead Enricher — mix websites and Facebook Pages in one list and get contacts for both in one run.
- Luma Events Scraper — list the hosts of tech and startup events, then pull contacts from their websites here.
- Meetup Groups Scraper — find local community organizers, then turn their group websites into emails and phone numbers.