Contact Details Scraper - Emails, Phones & Socials avatar

Contact Details Scraper - Emails, Phones & Socials

Pricing

$1.50 / 1,000 websites

Go to Apify Store
Contact Details Scraper - Emails, Phones & Socials

Contact Details Scraper - Emails, Phones & Socials

Get emails, phone numbers and addresses from a list of websites. Contact details are rarely on the home page. It checks /contact, /about and the German /impressum page. You also get contact forms and 10 social profiles. Emails written with [at] are read too. $1.50 per 1,000 sites, flat.

Pricing

$1.50 / 1,000 websites

Rating

0.0

(0)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

6

Total users

2

Monthly active users

a day ago

Last modified

Share

Contact Details Scraper: emails, phones, addresses and socials from a list of websites

Paste a list of domains and get one row per site holding whatever contact details that site publishes in public. Email addresses, phone numbers, postal addresses, social profiles, and the contact forms themselves, down to the field names each form expects. Every finding says which page it came from.

It does not stop at the home page, because contact details rarely live there. Each site gets a short crawl weighted towards contact, about, team, legal and the German imprint page. And plenty of businesses publish no email address at all any more, so a row with social links and a form on it, and emails: [], is a normal result rather than a failure.

InputDomains or website URLs, one per line
OutputOne row per website
Ceiling20 pages crawled per site
Account neededNone
Price$1.50 per 1,000 websites, flat on every plan

🔍 What Contact Details Scraper does

For each site it fetches the home page, picks the links most likely to carry contact details, and reads up to maxPagesPerSite of them. Everything it finds is merged into one record for that domain, deduplicated, with the source page recorded against each item.

Obfuscated addresses are decoded, so hello [at] example [dot] com comes back as a real address. An email that only exists as an image does not, because that would need character recognition, and it is the one dodge that still works against this.

Ten social platforms are recognised by name: LinkedIn, Facebook, Instagram, X, YouTube, TikTok, Pinterest, GitHub, Threads and Bluesky. Anything else stays a plain link and is not classified.

📥 What you give it

{
"websites": ["apify.com", "https://stripe.com", "example.co.uk"],
"maxPagesPerSite": 8,
"defaultCountry": "US"
}
FieldDefaultWhat it is
websitesnoneDomains or full URLs, one per line. A value with no scheme is treated as HTTPS. The Console box starts with one example; an API call has to send its own.
startUrlsnoneThe Apify request-list format, if that is what you already have. Combine it with websites freely.
maxPagesPerSite61 to 20. More pages finds more and takes longer, and it does not change what you pay.
defaultCountrynoneA two-letter code like US, GB or CA, used to make sense of phone numbers written without a country code.
maxConcurrency51 to 20 sites at once. Each site is still crawled one page at a time.
respectRobotsTxtonSkips pages the site asks crawlers not to read.
requestTimeoutSecs123 to 45 seconds per page.
maxResponseSizeKb1536128 to 5120. A page bigger than this is skipped.
proxyConfigurationoffOptional network settings, off because most business sites answer a direct request fine.

startUrls wants actual URL entries. A requestsFromUrl pointing at a text file of links is not picked up, so paste those URLs in rather than linking them.

📤 What you get back

A real row from a recent run, with the repeated source-page lists cut short:

{
"ok": true,
"siteUrl": "https://www.python.org/about/quotes/",
"requestedUrl": "https://www.python.org/",
"domain": "python.org",
"title": "Welcome to Python.org",
"description": "The official home of the Python Programming Language",
"emails": [],
"phones": [],
"socialLinks": [
{ "platform": "linkedin", "url": "https://www.linkedin.com/company/python-software-foundation/",
"sourcePages": ["https://www.python.org/", "..."] },
{ "platform": "x", "url": "https://twitter.com/ThePSF", "sourcePages": ["..."] },
{ "platform": "github", "url": "https://github.com/python/cpython/issues", "sourcePages": ["..."] }
],
"addresses": [],
"contactForms": [],
"jsonLdOrganizations": [],
"jsonLdPeople": [],
"crawledPages": [
{ "url": "https://www.python.org/", "status": 200, "title": "Welcome to Python.org", "contactsFound": 4 }
],
"pagesCrawled": 4,
"pagesAttempted": 4,
"emailCount": 0,
"phoneCount": 0,
"socialLinkCount": 5,
"robotsTxtRespected": true,
"pageErrors": []
}
FieldWhat it is
sourcePagesOn every finding. When someone asks where an address came from, you point at the page instead of checking it again by hand.
emailsEach entry is the address plus the pages it appeared on.
phonesThe number as written, plus e164 and country once it has been normalised. That normalisation leans on defaultCountry for locally written numbers.
contactFormsThe form's action URL, method and field names, and whether it has message, email and phone fields. Worth reading when a site publishes no address at all.
jsonLdOrganizations, jsonLdPeoplePulled from the site's own structured data. Named people with job titles turn up here more often than you would expect.
crawledPagesOne entry per page fetched, with its status, title and how many findings came off it.
pageErrorsPages that could not be read, each with its own code, such as a page the site asked crawlers to skip.
domainThe site key, without www. Two inputs that resolve to the same site collapse into one row.

The preview table shows the counts rather than the addresses themselves. Download the JSON or CSV to get the actual emails, phones and links.

🧾 Reading the output

Three kinds of row can land in your dataset.

RowHow to spot itCharged
A siteok: true and a domainyes
The sample row_sample: trueno
A diagnosticok: false and an errorCodeno
CodeWhat it means
BAD_INPUTNothing usable in websites or startUrls.
NO_CONTACTSThe site loaded fine and published nothing this actor recognises.
NO_RESULTSNothing on that site could be read. Often a domain that redirects somewhere else, or a site that asks crawlers to stay out.
UNSAFE_URLThe hostname resolves to a private address, so it was not fetched.
TIMEOUTThe site was too slow inside requestTimeoutSecs.
BLOCKEDThe site refused the request.
INTERNAL_ERRORSomething went wrong in the run itself. Send us the run ID.

One bad site never stops the others.

▶️ How to run it

  1. Open Contact Details Scraper and click Try for free.
  2. Paste your domains into Websites, one per line.
  3. Set Default phone country if your list is mostly one country. It makes the phone numbers usable.
  4. Raise Maximum pages per site if you want a deeper look, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$1.50 per 1,000 websites. Flat on every Apify plan, no volume tiers.

You pay per website, not per page. A site where 20 pages get read costs exactly the same as one where a single page does, so raising maxPagesPerSite buys you more findings for the same money. The only thing it costs is time.

Sites that failed to load, sites with nothing to find, pages skipped for size, and duplicate entries that collapsed into one row are all free, and a run where no site turned anything up costs you nothing.

💡 What people use it for

  • Turning a list of company domains into a contact sheet, with the source page recorded beside each finding so nothing has to be re-checked.
  • Finding the form when there is no email. The field names are what you need to reach a company programmatically.
  • Pulling the ten social profiles for a set of brands in one pass, instead of hunting footers.
  • Checking a supplier list still has working, published contact details on it.
  • Reading named people and roles out of a site's structured data.

🚧 What it does not do

  • No JavaScript. Details injected by a script after the page loads are not seen. Most business sites publish them in the HTML, but a footer fetched from an API is a genuine miss.
  • No logins. Public pages only, nothing behind a sign-in, a paywall or a form.
  • No emails inside images. That needs character recognition, which this does not do.
  • Same site only. Subdomains and a domain that redirects somewhere else are not followed.
  • 20 pages per site is the ceiling. This is a targeted crawl of the likely pages, not a full spider.
  • It respects robots.txt by default, so a site that asks crawlers to stay out returns little. That is deliberate.
  • It finds what is published, not what is true. A stale address on an old page still comes back.
  • No verification. It does not check whether an address accepts mail.

🧭 Which contact scraper do you need?

If you wantUse
Emails, phones and socials from a list of websitesThis one
Contact details off a Facebook PageFacebook Page Contact Info Scraper
The people at a companyLinkedIn Company Employees Scraper
Local businesses with phone and addressGoogle Maps Scraper
To check a list of addresses is deliverableEmail Verifier

❓ Questions people ask

Does reading more pages cost more? No. The charge is per website. Deeper crawls cost time, not money.

A site came back with no emails. Is that a bug? Usually not. A lot of companies publish a form instead of an address now. Check contactForms and socialLinks on that row.

Why is one of my phone numbers formatted oddly? A locally written number needs defaultCountry to be read correctly. Set it to match your list.

Can it read a page that needs a login? No. It sees what any visitor sees.

Can I schedule it? Yes, like any Apify actor. A saved list of domains makes it repeatable.

Is collecting contact details legal? These are details a business chose to publish. They are still personal data under GDPR and similar laws, and marketing to them has its own rules, so have a lawful basis and honour opt-outs. Apify's write-up on scraping and the law is a good place to start, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the website and the run ID. The errorCode on the diagnostic row, and pageErrors on a real row, usually name the problem between them.