Website Email & Contact Scraper - Find Emails, Phones & Socials
Pricing
from $4.00 / 1,000 results
Website Email & Contact Scraper - Find Emails, Phones & Socials
Feed any list of website URLs and get clean CSV/JSON of public emails, phone numbers, and social links from each contact page. Skip the $1,250/mo ZoomInfo bills and 50% bounce rates
Pricing
from $4.00 / 1,000 results
Rating
0.0
(0)
Developer
Anas Nadeem
Maintained by CommunityActor stats
0
Bookmarked
98
Total users
8
Monthly active users
21 days ago
Last modified
Categories
Share
Website Email & Contact Scraper
Extract public email addresses, phone numbers from telephone links, social links, and page URLs from websites. Export results as CSV or JSON with source URLs for reviewing and enriching company records.
Use this Actor to collect website contact information, prepare CRM imports, or inventory links across a small set of pages. It extracts published information; it does not verify email deliverability or guarantee that a site contains contact details.
Quick start
- Add one or more website URLs under Start URLs.
- Choose Full page-wise output for contacts and links, or Emails-only output for individual email records.
- Start with HTTP mode, 5 pages, depth 1, and concurrency 1.
- Run the Actor and inspect Output. Review the crawl summary as well as the contact records.
- Export CSV or JSON. Filter by
recordTypebefore importing contacts into another tool.
Paste this into the JSON input editor for a small TrustMRR website example:
{"startUrls": [{ "url": "https://trustmrr.com/" }],"mode": "full","crawlerType": "http","maxPages": 5,"maxDepth": 1,"sameDomainOnly": true,"includeSubdomains": false,"extractEmails": true,"extractPhones": true,"extractSocial": true,"extractLinks": true,"extractImages": false,"extractFiles": false,"maxConcurrency": 1,"requestTimeoutSecs": 30}
This example examines TrustMRR's own pages. It does not visit the websites of listed startups or extract their revenue data. Live content can change, and contact arrays can be empty. A complete saved input is available in examples/trustmrr-task-input.json.
Inputs and crawl scope
| Input | Default | Behavior |
|---|---|---|
startUrls | Required | Website URLs to begin crawling. |
mode | full | Page records or individual email records. |
crawlerType | http | HTTP reads returned HTML. See browser limitation below. |
maxPages | 200 | Shared page budget for the whole run, not per website. Concurrency can cause slight overshoot. |
maxDepth | 2 | Seeds are depth 0; directly linked pages are depth 1. |
sameDomainOnly | true | Restricts discovered links to the seed hostname. |
includeSubdomains | false | Also allows subdomains of the seed hostname. |
includeUrlGlobs | [] | Only follow discovered URLs matching an included pattern when set. |
excludeUrlGlobs | [] | Skip matching discovered URLs. |
extractEmails, extractPhones, extractSocial | true | Enable contact extraction. Email-only mode always extracts emails. |
extractLinks, extractImages, extractFiles | true | Include page, image, and downloadable file URLs. Files are not downloaded. |
maxConcurrency | 25 | Parallel requests; use 1 for a small demonstration. |
requestTimeoutSecs | 20 | Request-handler timeout; not a total run duration limit. |
proxyConfiguration | Direct connection | Optional Apify proxy configuration. |
Include/exclude patterns apply to discovered links, not the initial seed URLs. For a contact-focused crawl, try includeUrlGlobs: ["**/contact**", "**/about**", "**/team**"]. This follows matching links already present on a page; it does not guess missing contact URLs. Use the site's final canonical hostname to avoid apex/www redirect scope issues.
Output and exports
All records are stored in the default dataset. The Actor output schema links to this dataset; dataset views organize its columns. The Overview view includes fields for both modes and summaries. The Emails view selects email-related columns but does not remove summary rows.
recordType | When produced | Contents |
|---|---|---|
page | Full mode | seedUrl, pageUrl, title, depth, statusCode, contact/link arrays, and counts. |
email | Email-only mode | seedUrl, url, email, depth, and statusCode. |
seed_summary | Both modes | Per-seed pages crawled, failed requests, unique emails, status histogram, and totals. |
run_summary | Both modes | Overall pages, failures, unique emails, totals, and duration. |
Email-only output deduplicates emails within each seed. The same address can appear for different seeds. Full mode retains emails on each page where they are found. Aggregate counts can include repeated contacts; uniqueEmails reports distinct addresses within the relevant scope.
Illustrative email record (not a measured TrustMRR result):
{"recordType": "email","seedUrl": "https://company.test/","url": "https://company.test/contact","email": "hello@company.test","depth": 1,"statusCode": 200}
Full-mode contact fields are emails and phoneNumbers arrays, plus a socialLinks object grouped by platform. Link fields are internalLinks, externalLinks, images, and files. Disabled extraction fields return empty collections. JSON preserves nested structures; choose appropriate columns when exporting CSV.
Pricing
The configured base price is $0.004 per default dataset item ($4 per 1,000 items). Check the Actor's Pricing tab for the price applicable to your account.
The current implementation writes page records in full mode or email records in email-only mode, plus one summary per seed and one run summary. Include summary items when budgeting; a dataset row is not necessarily a contact. For example, five page records from one seed plus two summaries total seven items, or $0.028 at the base rate. This is an illustration, not a guaranteed run charge.
maxPages limits crawl work, not the number of email rows: a single page can contain many emails. Set a maximum charge in Console when available to control spend.
API and workflow integration
Run the Actor through the Apify API, then retrieve its default dataset. For email-only integrations, retain only rows where recordType equals email.
curl -X POST \"https://api.apify.com/v2/acts/whoareyouanas~website-link-scraper/run-sync-get-dataset-items" \-H "Authorization: Bearer YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls":[{"url":"https://trustmrr.com/"}],"mode":"emails_only","crawlerType":"http","maxPages":5,"maxDepth":1,"maxConcurrency":1}'
For longer crawls, start a run asynchronously, wait for completion, and fetch the dataset. In n8n or Make, connect the Actor run to a record-type filter, then map email and url into your destination. Add a separate email-verification step if deliverability matters.
Current limitations
- Phones are extracted from
tel:links only. Plain-text phone numbers are not currently extracted. - HTTP mode cannot read contacts that appear only after client-side JavaScript runs. The browser option requires Playwright and Chromium in the deployed image; the current repository Docker setup has not been validated for this mode. Use HTTP for the example task.
- Email extraction supports basic text obfuscation such as
name [at] company [dot] com, but can miss protected addresses or return false positives. It does not verify ownership or deliverability. - Social links are classified by platform with some share-link filtering; not every returned URL is a company profile.
- Requests can fail or return no contacts. Review
failedRequests,pagesCrawled, and contact arrays; a successful run status alone does not establish complete coverage. - Multiple seeds share a URL-deduplicated request queue and a global page budget. Overlapping seeds may not receive separate page coverage.
Feedback
Report issues with your input, run URL, expected result, and affected page. Remove secrets before sharing input. If the Actor is useful, an honest Store review helps other users evaluate it.
Local development
npm installnpm run buildnpm test
Provide local input through Apify storage before running npm run dev. Actor configuration and schemas live in .actor/.