Website Email & Contact Scraper - Find Emails, Phones & Socials avatar

Website Email & Contact Scraper - Find Emails, Phones & Socials

Pricing

from $4.00 / 1,000 results

Go to Apify Store
Website Email & Contact Scraper - Find Emails, Phones & Socials

Website Email & Contact Scraper - Find Emails, Phones & Socials

Feed any list of website URLs and get clean CSV/JSON of public emails, phone numbers, and social links from each contact page. Skip the $1,250/mo ZoomInfo bills and 50% bounce rates

Pricing

from $4.00 / 1,000 results

Rating

0.0

(0)

Developer

Anas Nadeem

Anas Nadeem

Maintained by Community

Actor stats

0

Bookmarked

98

Total users

8

Monthly active users

21 days ago

Last modified

Share

Website Email & Contact Scraper

Extract public email addresses, phone numbers from telephone links, social links, and page URLs from websites. Export results as CSV or JSON with source URLs for reviewing and enriching company records.

Use this Actor to collect website contact information, prepare CRM imports, or inventory links across a small set of pages. It extracts published information; it does not verify email deliverability or guarantee that a site contains contact details.

Quick start

  1. Add one or more website URLs under Start URLs.
  2. Choose Full page-wise output for contacts and links, or Emails-only output for individual email records.
  3. Start with HTTP mode, 5 pages, depth 1, and concurrency 1.
  4. Run the Actor and inspect Output. Review the crawl summary as well as the contact records.
  5. Export CSV or JSON. Filter by recordType before importing contacts into another tool.

Paste this into the JSON input editor for a small TrustMRR website example:

{
"startUrls": [{ "url": "https://trustmrr.com/" }],
"mode": "full",
"crawlerType": "http",
"maxPages": 5,
"maxDepth": 1,
"sameDomainOnly": true,
"includeSubdomains": false,
"extractEmails": true,
"extractPhones": true,
"extractSocial": true,
"extractLinks": true,
"extractImages": false,
"extractFiles": false,
"maxConcurrency": 1,
"requestTimeoutSecs": 30
}

This example examines TrustMRR's own pages. It does not visit the websites of listed startups or extract their revenue data. Live content can change, and contact arrays can be empty. A complete saved input is available in examples/trustmrr-task-input.json.

Inputs and crawl scope

InputDefaultBehavior
startUrlsRequiredWebsite URLs to begin crawling.
modefullPage records or individual email records.
crawlerTypehttpHTTP reads returned HTML. See browser limitation below.
maxPages200Shared page budget for the whole run, not per website. Concurrency can cause slight overshoot.
maxDepth2Seeds are depth 0; directly linked pages are depth 1.
sameDomainOnlytrueRestricts discovered links to the seed hostname.
includeSubdomainsfalseAlso allows subdomains of the seed hostname.
includeUrlGlobs[]Only follow discovered URLs matching an included pattern when set.
excludeUrlGlobs[]Skip matching discovered URLs.
extractEmails, extractPhones, extractSocialtrueEnable contact extraction. Email-only mode always extracts emails.
extractLinks, extractImages, extractFilestrueInclude page, image, and downloadable file URLs. Files are not downloaded.
maxConcurrency25Parallel requests; use 1 for a small demonstration.
requestTimeoutSecs20Request-handler timeout; not a total run duration limit.
proxyConfigurationDirect connectionOptional Apify proxy configuration.

Include/exclude patterns apply to discovered links, not the initial seed URLs. For a contact-focused crawl, try includeUrlGlobs: ["**/contact**", "**/about**", "**/team**"]. This follows matching links already present on a page; it does not guess missing contact URLs. Use the site's final canonical hostname to avoid apex/www redirect scope issues.

Output and exports

All records are stored in the default dataset. The Actor output schema links to this dataset; dataset views organize its columns. The Overview view includes fields for both modes and summaries. The Emails view selects email-related columns but does not remove summary rows.

recordTypeWhen producedContents
pageFull modeseedUrl, pageUrl, title, depth, statusCode, contact/link arrays, and counts.
emailEmail-only modeseedUrl, url, email, depth, and statusCode.
seed_summaryBoth modesPer-seed pages crawled, failed requests, unique emails, status histogram, and totals.
run_summaryBoth modesOverall pages, failures, unique emails, totals, and duration.

Email-only output deduplicates emails within each seed. The same address can appear for different seeds. Full mode retains emails on each page where they are found. Aggregate counts can include repeated contacts; uniqueEmails reports distinct addresses within the relevant scope.

Illustrative email record (not a measured TrustMRR result):

{
"recordType": "email",
"seedUrl": "https://company.test/",
"url": "https://company.test/contact",
"email": "hello@company.test",
"depth": 1,
"statusCode": 200
}

Full-mode contact fields are emails and phoneNumbers arrays, plus a socialLinks object grouped by platform. Link fields are internalLinks, externalLinks, images, and files. Disabled extraction fields return empty collections. JSON preserves nested structures; choose appropriate columns when exporting CSV.

Pricing

The configured base price is $0.004 per default dataset item ($4 per 1,000 items). Check the Actor's Pricing tab for the price applicable to your account.

The current implementation writes page records in full mode or email records in email-only mode, plus one summary per seed and one run summary. Include summary items when budgeting; a dataset row is not necessarily a contact. For example, five page records from one seed plus two summaries total seven items, or $0.028 at the base rate. This is an illustration, not a guaranteed run charge.

maxPages limits crawl work, not the number of email rows: a single page can contain many emails. Set a maximum charge in Console when available to control spend.

API and workflow integration

Run the Actor through the Apify API, then retrieve its default dataset. For email-only integrations, retain only rows where recordType equals email.

curl -X POST \
"https://api.apify.com/v2/acts/whoareyouanas~website-link-scraper/run-sync-get-dataset-items" \
-H "Authorization: Bearer YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://trustmrr.com/"}],"mode":"emails_only","crawlerType":"http","maxPages":5,"maxDepth":1,"maxConcurrency":1}'

For longer crawls, start a run asynchronously, wait for completion, and fetch the dataset. In n8n or Make, connect the Actor run to a record-type filter, then map email and url into your destination. Add a separate email-verification step if deliverability matters.

Current limitations

  • Phones are extracted from tel: links only. Plain-text phone numbers are not currently extracted.
  • HTTP mode cannot read contacts that appear only after client-side JavaScript runs. The browser option requires Playwright and Chromium in the deployed image; the current repository Docker setup has not been validated for this mode. Use HTTP for the example task.
  • Email extraction supports basic text obfuscation such as name [at] company [dot] com, but can miss protected addresses or return false positives. It does not verify ownership or deliverability.
  • Social links are classified by platform with some share-link filtering; not every returned URL is a company profile.
  • Requests can fail or return no contacts. Review failedRequests, pagesCrawled, and contact arrays; a successful run status alone does not establish complete coverage.
  • Multiple seeds share a URL-deduplicated request queue and a global page budget. Overlapping seeds may not receive separate page coverage.

Feedback

Report issues with your input, run URL, expected result, and affected page. Remove secrets before sharing input. If the Actor is useful, an honest Store review helps other users evaluate it.

Local development

npm install
npm run build
npm test

Provide local input through Apify storage before running npm run dev. Actor configuration and schemas live in .actor/.