Fast Website Contact Extractor
Pricing
from $0.50 / 1,000 results
Fast Website Contact Extractor
Extract public emails, phone numbers, social profiles, and contact pages from up to five HTML pages.
Pricing
from $0.50 / 1,000 results
Rating
0.0
(0)
Developer
microautomation lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Extract public contact information from one website with a small, predictable HTTP-only crawl. Each successful run writes one record to the Default Dataset.
What it extracts
- Email addresses, including
mailto:links. - Phone numbers, including
tel:links. - Public Facebook, Instagram, LinkedIn, X/Twitter, YouTube, TikTok, Pinterest, and Threads profile URLs.
- Same-site Contact/About-style page URLs actually found in page links.
- The exact source page URL for every email, phone number, and social link.
Duplicate contact values and duplicate page URLs are removed. The first page on which a value is found is retained as its sourceUrl.
Input
Provide exactly one absolute HTTP or HTTPS URL:
{ "url": "https://example.com/" }
The Actor is for public websites only. 1 run = 1 website; bulk URL input is not supported. URLs containing credentials are rejected.
How it works
The Actor fetches the input page using normal HTTP, discovers same-origin Contact/About-style links present in the HTML, and checks those links in discovery order. It checks a maximum of 5 HTML pages total, never intentionally crawls external domains, never guesses an unlimited list of paths, and never fetches the same scheduled URL twice.
This is HTTP-only extraction. JavaScript is not executed. It is not a full-site crawler and does not use browser automation, Playwright, Puppeteer, proxies, AI/LLM, search engines, login, paid APIs, a database, or CRM integrations. It performs no email or SMTP verification. Extracted information comes only from publicly accessible pages returned by the website.
The whole run has a 15-second timeout, including response-body reads and the additional discovered pages.
Output
{"url": "https://example.com/","emails": [{ "value": "hello@example.com", "sourceUrl": "https://example.com/contact" }],"phones": [{ "value": "+1 212 555 0100", "sourceUrl": "https://example.com/contact" }],"socialLinks": [{ "url": "https://linkedin.com/company/example", "sourceUrl": "https://example.com/" }],"contactPages": ["https://example.com/contact"],"sourcePages": ["https://example.com/", "https://example.com/contact"],"pagesChecked": 2,"httpStatus": 200,"finalUrl": "https://example.com/"}
httpStatus and finalUrl describe the top-page response. sourcePages contains successfully processed HTML pages. An empty contact result is a valid success and uses empty arrays.
Pricing
This Actor uses pay-per-event pricing: $0.00005 for each Actor Start and $0.00050 for the result event. A successful run produces one dataset record and charges the result event once. Failed runs do not produce a result record. Volume discounts are currently off.
Errors
Input or fetch failures end the run with exit code 1 and do not write incomplete success data to the Default Dataset.
| Code | Meaning |
|---|---|
INVALID_URL | Missing or invalid absolute HTTP(S) input URL, or a URL containing credentials. |
HTTP_ERROR | A requested page returned a non-2xx response. |
NETWORK_ERROR | DNS, connection, TLS, or another fetch failure. |
TIMEOUT | The run exceeded 15 seconds. |
NON_HTML | A requested response was not HTML or XHTML. |
Local development
Requires Node.js 22 or later and the Apify CLI.
npm installnpm testnpm run typechecknpm run buildapify run --no-purge --input-file input.json
Do not use a production website for deterministic tests; the automated test suite uses local HTML fixtures and a local HTTP server.