# Changelog of Website Contact Details Crawler (`yasaslive/website-contact-details-crawler`) Actor

- **URL**: https://apify.com/yasaslive/website-contact-details-crawler/changelog.md
- **Full Actor documentation**: https://apify.com/yasaslive/website-contact-details-crawler.md

## Changelog

All notable changes to this Actor are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the project uses [Semantic Versioning](https://semver.org/).

### \[1.0.0] — 2026-09-26

#### Added

- HTTP-only `CheerioCrawler` that crawls each website (same domain by default) and returns one deduplicated record per domain: emails, phones (E.164 + raw + country), social profiles, organisation name, address and a per-value `sources` map.
- Email extraction from `mailto:` links (entity- and percent-encoded), visible text, `meta`/`title`/`aria-label` attributes, Cloudflare `data-cfemail` / `/cdn-cgi/l/email-protection#…` decoding, and `[at]`/`(dot)`/`name at domain dot com` obfuscation.
- Phone extraction with `libphonenumber-js` (full "max" metadata for accurate validation) from `tel:` links, JSON-LD and formatted numbers in text, using `phoneCountryHint` for national formats.
- Social profiles for Facebook, Instagram, LinkedIn company + personal, X/Twitter, YouTube, TikTok, Pinterest, GitHub, Threads and Bluesky, canonicalised and with share widgets / posts filtered out.
- `schema.org` Organization / LocalBusiness (and ~60 subtypes) JSON-LD parsing, including `@graph`, `contactPoint`, `sameAs` and lenient parsing of slightly malformed blocks.
- Pay-per-event charging: `actor-start`, `page-crawled`, `domain-with-contacts` (see `.actor/pay_per_event.json`).
- `STATS` (logged and persisted every 30 s) with categorised error counters, and `FAILED_DOMAINS` in the default key-value store.
- Migration-safe crawl state via `Actor.useState`.
- Unit tests on saved fixtures and an end-to-end test that runs the Actor against a local fake web with simulated PPE pricing.

#### Design decisions

- **Runtime is Node 22, not 20.** The spec asked for the `apify/actor-node:20` base, but Node 20 reached end-of-life in April 2026 and gets no security fixes. Node 22 still meets "Node.js 20+".
- **robots.txt respected by default** (`respectRobotsTxt: true`, user-toggleable). The owner chose this for Store/ToS safety. Crawlee's built-in `respectRobotsTxtFile` is used; skipped pages are counted in `STATS.pagesSkippedRobotsTxt`.
- **Websites without contacts are still delivered** (`hasContacts: false`) because their `page-crawled` events were already paid for. Only the `domain-with-contacts` value event depends on finding an email or phone. Websites where **no** page loaded are not pushed (nothing was charged for them); they are listed in `FAILED_DOMAINS` instead. This follows the "never emit an item you did not charge for" rule: every dataset row is backed by at least one charged `page-crawled` event.
- **`maxResults` counts delivered websites**, with or without contacts. Websites are fed into the queue in a window (at most `maxResults − delivered` and 25 at a time), so the run never crawls more websites than it can deliver.
- **Budget reserve.** Before charging a page, the Actor keeps enough budget for a `domain-with-contacts` event for every website in progress. Users therefore never pay for pages of a website whose record could not be delivered. When the budget can't cover another website, it stops starting new ones and finishes the ones in flight. When it can't cover another page, the crawler stops. All charges go through a single mutex.
- **Local PPE simulation.** Locally the SDK's `calculateMaxEventChargeCountWithinLimit` prices every event at $1. The Actor computes the same formula from the configured prices when not on the platform, and on the platform takes the minimum of both.
- **Priority crawl.** Links matching `prioritizePaths` (whole path segment, so `/en/contact-us/` matches) go to the front of the queue. Literal priority paths (the first 3) are only *guessed* when the homepage links to none of them, which avoids a burst of 404s. 404s are free.
- **Politeness.** `sameDomainDelaySecs: 1` means at most about one request per second per website, while many websites are crawled in parallel.
- **Soft-block detection** only uses markers specific to challenge interstitials (Cloudflare `_cf_chl_opt`, "Just a moment…", DataDome, PerimeterX, Incapsula, SiteGround). It never keys on the word "captcha" alone, because contact forms often embed reCAPTCHA. A redirect from a public page onto a login page is treated as "login required" and is not retried. HTTP 401 is handled the same way.
- **Retries** use exponential backoff with jitter (1 s, 2 s, 4 s … capped at 30 s) and rotate the session/IP on blocks, rate limits and proxy errors.
- **Text phone matches must look formatted** (spaces, dashes, dots, brackets or `+`). This removes order numbers, SKUs and prices. Numbers in `tel:` links and JSON-LD are accepted unformatted.
- **Start URLs** accept bare domains, full URLs and remote lists (`requestsFromUrl`). Private, link-local and internal hostnames are rejected as an SSRF guard. If a start page fails over HTTPS with a network error, it is retried once over HTTP.
- **Demo input**: russanddaughters.com, levainbakery.com, voodoodoughnut.com. These are public small-business sites whose robots.txt allows crawling. They were vetted on 26 Sep 2026: each returned emails and phones over plain HTTP. The rejected candidates were tartinebakery.com and veniero.com (no contacts in server HTML), joespizzanyc.com (no email), and zingermansdeli.com and katzsdelicatessen.com (bot-blocked).

#### Known issues / accepted risks

- `npm audit` reports 13 *moderate* advisories. They all trace to one transitive dependency, `stream-json@1.9.1` via `@crawlee/core` (GHSA-528h-pc64-c93x, a DoS in the `pick`/`ignore`/`filter`/`replace` filters). Crawlee only uses `StreamArray` to deserialise its own persisted state, not the vulnerable filters, and never on attacker-controlled input. The fix requires a major version bump that Crawlee has not adopted yet. Re-check when upgrading Crawlee.
