# Changelog of B2B Lead Enrichment (`teodor_banea/b2b-lead-enrichment-free`) Actor

- **URL**: https://apify.com/teodor_banea/b2b-lead-enrichment-free/changelog.md
- **Full Actor documentation**: https://apify.com/teodor_banea/b2b-lead-enrichment-free.md

## Changelog

### 2.1.0 - 2026-10-02

Reliability, cost and data-quality release. Output fields are unchanged; values are better.

#### Reliability

- **Large runs no longer run out of memory.** A 40-company run on 1 GB was OOM-killed
  mid-run in 2.0, losing every company not yet written. Pages are now stripped of inline
  scripts, styles and SVG before parsing (88 MB → 8 MB of DOM for five site-builder pages),
  and the worker pool stops admitting companies while container memory is high.
- **Concurrency scales to the container.** The pool starts with two companies and adds one
  per second while the CPU keeps up, backing off on platform CPU-overload signals or
  event-loop lag, instead of starting ten at once on a quarter of a CPU core.
- **Every page fetch has a hard deadline.** Some hosts never answered and ignored the HTTP
  client's own timeout, holding the company until its budget ran out and discarding
  everything else found. Such hosts are now read through the browser.
- **Slow sites return partial data.** The website scraper stops a little before its
  deadline and keeps the pages it has, instead of being abandoned with nothing.
- The browser fallback uses the Chromium preinstalled in Apify's base image, so it works
  regardless of the Playwright revision.

#### Cost

- The image adds ~50 MB to Apify's cached base image instead of a second Chromium and dev
  tooling (~1.2 GB compressed in all), which every run on a cold worker had to pull first.
- Crawlee's request queue is gone (companies are already in memory): no request-queue
  writes, no periodic statistics or session-pool snapshots, and a faster start.
- The providers for a company run in parallel instead of one after another.
- Playwright is loaded only when a page actually needs a browser.
- A user's maximum cost per run is respected before work is done, not just when rows are
  written: companies past the limit are no longer enriched only to be discarded.

#### Data quality

- **Named decision-makers from team, about and leadership pages**, with job title and
  LinkedIn profile, even when no email is published. Testimonials and reviews are excluded.
  When the company publishes some personal addresses, their format is applied to other
  named staff (`personEmails[].type: "pattern"`, with a confidence).
- **The primary contact is always reachable**: the best contact with an email address.
- **Bot walls and parked domains are recognised.** Cloudflare challenge and block pages are
  no longer scraped into a company called "Attention Required!" with an IP address for a
  phone number; parked, expired and for-sale domains are not billed.
- **Company names** are chosen by matching every candidate (schema.org, og:site_name, each
  title segment, logo alt text, copyright holder) against the domain, including the domain
  a site redirects to. Taglines and page names no longer become the company name, and
  registered names ("… SRL", "… LLP") go to `companyLegalName`.
- **Locations**: street addresses no longer land in the city field, `<br>`-separated address
  blocks are parsed, and the head office is preferred when several offices are listed.
- **schema.org** LocalBusiness subtypes (Dentist, AccountingService, Winery…) are read for
  contact data and used to classify the industry; `email` and `telephone` are used.
- **Industry** classification covers far more small-business verticals and matches whole
  words (no more "cpa" inside "cpanel").
- **Removed a fabricated headcount**: a "100-500" range was inferred from the company's
  email-security vendor and was wrong as often as right. `companySizeEmployeesRange` now
  only carries a range the company publishes.
- **GitHub** organisations are attached only when their listed website matches the domain.
- Social links are validated by hostname; dates and IP addresses are no longer phones;
  questions and prose are no longer job titles; many more multilingual and functional
  shared inboxes are recognised; duplicate contacts are merged; empty strings are `null`.

### 2.0.2 - 2026-09-13

Platform hardening. Measured on 30 companies: 30/30 succeeded, 0 failed, 156s, 0.043
compute units - roughly a third of the cost per result the Actor used to run at.

- **Dataset writes are serialised.** Called concurrently from every worker,
  `Actor.pushData` stopped settling for the later companies. The request handler waited on
  a write that never returned, was killed by Crawlee at its timeout, and a company that had
  been enriched successfully was lost. Writes now queue one at a time, are time-boxed, and
  are retried once against a fresh handle. A company whose row cannot be stored is reported
  rather than dropped, and is never counted as billable.
- **A company is bounded by one deadline** shared across all providers, each receiving only
  the time that remains. Previously each provider had its own timeout and the worst case
  summed to more than the request handler allowed, so a slow company was killed and lost
  everything it had gathered instead of finishing with partial data.
- **Companies are never retried.** Providers already retry their own requests; re-running a
  whole enrichment costs twice and risked charging twice.
- **Redundant DOM parsing removed.** Each page was parsed about four times, including a
  full serialise-and-reparse per call in `visibleText`. Now one parse per page.
- Chromium is installed into a writable path in the image, so the browser fallback
  actually works on the platform instead of silently degrading to HTTP-only.
- Status messages are throttled and never awaited, and named-dataset access is time-boxed
  and self-disabling. Both could previously stall a company for minutes through client
  retry backoff.
- Domains are checked for an A or MX record before any page is fetched, so a dead domain
  no longer costs four browser page loads.

### 2.0.1 — 2026-09-13

Fixes for three failures that only appeared once the Actor ran on the Apify platform.

- **Chromium was missing from the image.** `playwright` resolved to 1.58.2, which looks
  for a browser revision the base image does not ship, so every browser fallback failed
  with "Executable doesn't exist at /pw-browsers/...". The browser is now installed
  explicitly during the build. Before this fix the Actor silently degraded to HTTP-only
  and returned no data for JavaScript-rendered sites.
- **Writing to the auxiliary `errors` dataset could stall a run.** The call returned
  "Insufficient permissions" and the Apify client retried it with exponential backoff for
  roughly two minutes, pushing the request handler past its 120-second timeout. Crawlee
  then retried the request and the company was enriched and pushed a second time, billing
  one input row twice. Named-dataset access is now time-boxed, and the first failure
  disables it for the rest of the run and falls back to logging.
- **A retried request can no longer re-push a company that was already pushed**, whatever
  the cause of the retry.

### 2.0.0 — 2026-09-13

Cost and data-quality rework, and a move to pay-per-result pricing.

#### Pricing

- The Actor is now billed **once per company enriched**. One input row produces exactly
  one dataset row regardless of how many contacts are found, so run cost is predictable
  from the CSV alone.
- Renamed from "B2B Lead Enrichment (Free)" to "B2B Lead Enrichment".

#### Breaking changes

- **Output is now one row per company** instead of one row per contact. Company fields are
  no longer duplicated across rows.
- `person*` columns are replaced by `primaryContact*` columns carrying the best contact
  found. The full list moved to the `contacts` array with a `contactCount` alongside it.
- `techStack` no longer contains SSL issuers, GitHub repository topics or source
  languages. These moved to `sslIssuer`, `githubTopics` and `programmingLanguages`.
- The mail host moved out of `techStack` into `emailProvider`.

#### Cost

- Pages are fetched over plain HTTP first; a browser is launched only for pages that
  render entirely in JavaScript. Previously every page load ran through Chromium.
- Removed a fixed 1.5-second wait on every page load, replaced by waiting for content to
  actually appear.
- Chromium is now launched lazily and shared, so runs that never meet a JavaScript-only
  site never start a browser at all.
- Browsers block images, fonts, media and stylesheets — only the DOM is ever read.
- Default concurrency raised from 3 to 10.
- Provider time budgets are now enforced by the waterfall, so one hung request can no
  longer hold a company open until the request handler times out.

Measured on a 12-company sample, per-company wall time fell from 2.55 s to 0.74 s and
browser page loads fell from 100% of fetches to roughly 10%.

#### Data quality

- **Sub-pages are discovered from the site's own navigation** instead of guessed from a
  fixed path list. The previous defaults never reached the contact page at all.
- **Documentation placeholders are rejected.** Addresses and phone numbers inside code
  samples are excluded, along with known example identities. Previously a Stripe API
  sample address and the placeholder number `1000000000` were published as real contacts.
- **Shared inboxes are surfaced** as labelled contacts and in `companyEmails`, ranked by
  sales relevance. They were previously collected and then discarded.
- **Contacts now carry names, titles, seniority and department** read from team and
  contact page cards, or derived from unambiguous `first.last@` addresses.
- **Locations are parsed rather than passed through.** Office lists, countries in the city
  field, combined "City, ST" values and non-places like "Remote" are handled correctly.
- Added `emailProvider`, `emailSecurity`, `emailMarketingTools`, `dnsProvider`,
  `sslIssuer`, `sslValidFrom`, `sslValidTo`, `companyLogoUrl`, `companyGithubUrl`,
  `companyEmails`, `companyPhone`, `programmingLanguages` and `githubTopics`.
- Added `enrichmentScore` and `fieldsFilled` for filtering thin records.
- Industry classification reads page headings as well as the description, weighted lower
  so marketing copy cannot outvote it, and covers 18 non-technology sectors.
- Homepages are retried on the `www` host and over HTTP before being given up on.
- schema.org parsing now walks `@graph` structures and picks the richest organization node.
- DNS falls back to DNS-over-HTTPS when the system resolver is unreachable.

#### Fixed

- `"San Francisco, CA"` resolved to country Canada, because `CA` is both a US state code
  and an ISO country code.
- A short page carrying schema.org or a meta description was treated as a JavaScript
  shell and needlessly re-fetched in a browser.
- `contact-us@` style addresses were treated as individuals, and a nearby office heading
  could be published as that person's name.
- Hyphenated surnames were dropped when deriving names from email addresses.
- "Chief Executive Officer" mapped to no department.

#### Added

- Unit test suite covering email validity, contact extraction, name derivation, phone
  validation, location normalisation, JavaScript-shell detection and sub-page discovery.
