Crawl any list of websites and extract every email address, phone number (E.164) and social media profile, deduplicated into one clean record per domain. Fast HTTP-only crawler with pay-per-result pricing.
HTTP-only CheerioCrawler that crawls each website (same domain by default) and returns one deduplicated record per domain: emails, phones (E.164 + raw + country), social profiles, organisation name, address and a per-value sources map.
Email extraction from mailto: links (entity- and percent-encoded), visible text, meta/title/aria-label attributes, Cloudflare data-cfemail / /cdn-cgi/l/email-protection#… decoding, and [at]/(dot)/name at domain dot com obfuscation.
Phone extraction with libphonenumber-js (full "max" metadata for accurate validation) from tel: links, JSON-LD and formatted numbers in text, using phoneCountryHint for national formats.
Social profiles for Facebook, Instagram, LinkedIn company + personal, X/Twitter, YouTube, TikTok, Pinterest, GitHub, Threads and Bluesky, canonicalised and with share widgets / posts filtered out.
schema.org Organization / LocalBusiness (and ~60 subtypes) JSON-LD parsing, including @graph, contactPoint, sameAs and lenient parsing of slightly malformed blocks.
Pay-per-event charging: actor-start, page-crawled, domain-with-contacts (see .actor/pay_per_event.json).
STATS (logged and persisted every 30 s) with categorised error counters, and FAILED_DOMAINS in the default key-value store.
Migration-safe crawl state via Actor.useState.
Unit tests on saved fixtures and an end-to-end test that runs the Actor against a local fake web with simulated PPE pricing.
Design decisions
Runtime is Node 22, not 20. The spec asked for the apify/actor-node:20 base, but Node 20 reached end-of-life in April 2026 and gets no security fixes. Node 22 still meets "Node.js 20+".
robots.txt respected by default (respectRobotsTxt: true, user-toggleable). The owner chose this for Store/ToS safety. Crawlee's built-in respectRobotsTxtFile is used; skipped pages are counted in STATS.pagesSkippedRobotsTxt.
Websites without contacts are still delivered (hasContacts: false) because their page-crawled events were already paid for. Only the domain-with-contacts value event depends on finding an email or phone. Websites where no page loaded are not pushed (nothing was charged for them); they are listed in FAILED_DOMAINS instead. This follows the "never emit an item you did not charge for" rule: every dataset row is backed by at least one charged page-crawled event.
maxResults counts delivered websites, with or without contacts. Websites are fed into the queue in a window (at most maxResults − delivered and 25 at a time), so the run never crawls more websites than it can deliver.
Budget reserve. Before charging a page, the Actor keeps enough budget for a domain-with-contacts event for every website in progress. Users therefore never pay for pages of a website whose record could not be delivered. When the budget can't cover another website, it stops starting new ones and finishes the ones in flight. When it can't cover another page, the crawler stops. All charges go through a single mutex.
Local PPE simulation. Locally the SDK's calculateMaxEventChargeCountWithinLimit prices every event at $1. The Actor computes the same formula from the configured prices when not on the platform, and on the platform takes the minimum of both.
Priority crawl. Links matching prioritizePaths (whole path segment, so /en/contact-us/ matches) go to the front of the queue. Literal priority paths (the first 3) are only guessed when the homepage links to none of them, which avoids a burst of 404s. 404s are free.
Politeness.sameDomainDelaySecs: 1 means at most about one request per second per website, while many websites are crawled in parallel.
Soft-block detection only uses markers specific to challenge interstitials (Cloudflare _cf_chl_opt, "Just a moment…", DataDome, PerimeterX, Incapsula, SiteGround). It never keys on the word "captcha" alone, because contact forms often embed reCAPTCHA. A redirect from a public page onto a login page is treated as "login required" and is not retried. HTTP 401 is handled the same way.
Retries use exponential backoff with jitter (1 s, 2 s, 4 s … capped at 30 s) and rotate the session/IP on blocks, rate limits and proxy errors.
Text phone matches must look formatted (spaces, dashes, dots, brackets or +). This removes order numbers, SKUs and prices. Numbers in tel: links and JSON-LD are accepted unformatted.
Start URLs accept bare domains, full URLs and remote lists (requestsFromUrl). Private, link-local and internal hostnames are rejected as an SSRF guard. If a start page fails over HTTPS with a network error, it is retried once over HTTP.
Demo input: russanddaughters.com, levainbakery.com, voodoodoughnut.com. These are public small-business sites whose robots.txt allows crawling. They were vetted on 26 Sep 2026: each returned emails and phones over plain HTTP. The rejected candidates were tartinebakery.com and veniero.com (no contacts in server HTML), joespizzanyc.com (no email), and zingermansdeli.com and katzsdelicatessen.com (bot-blocked).
Known issues / accepted risks
npm audit reports 13 moderate advisories. They all trace to one transitive dependency, stream-json@1.9.1 via @crawlee/core (GHSA-528h-pc64-c93x, a DoS in the pick/ignore/filter/replace filters). Crawlee only uses StreamArray to deserialise its own persisted state, not the vulnerable filters, and never on attacker-controlled input. The fix requires a major version bump that Crawlee has not adopted yet. Re-check when upgrading Crawlee.