Reliability, cost and data-quality release. Output fields are unchanged; values are better.
- Large runs no longer run out of memory. A 40-company run on 1 GB was OOM-killed
mid-run in 2.0, losing every company not yet written. Pages are now stripped of inline
scripts, styles and SVG before parsing (88 MB → 8 MB of DOM for five site-builder pages),
and the worker pool stops admitting companies while container memory is high.
- Concurrency scales to the container. The pool starts with two companies and adds one
per second while the CPU keeps up, backing off on platform CPU-overload signals or
event-loop lag, instead of starting ten at once on a quarter of a CPU core.
- Every page fetch has a hard deadline. Some hosts never answered and ignored the HTTP
client's own timeout, holding the company until its budget ran out and discarding
everything else found. Such hosts are now read through the browser.
- Slow sites return partial data. The website scraper stops a little before its
deadline and keeps the pages it has, instead of being abandoned with nothing.
- The browser fallback uses the Chromium preinstalled in Apify's base image, so it works
regardless of the Playwright revision.
- The image adds ~50 MB to Apify's cached base image instead of a second Chromium and dev
tooling (~1.2 GB compressed in all), which every run on a cold worker had to pull first.
- Crawlee's request queue is gone (companies are already in memory): no request-queue
writes, no periodic statistics or session-pool snapshots, and a faster start.
- The providers for a company run in parallel instead of one after another.
- Playwright is loaded only when a page actually needs a browser.
- A user's maximum cost per run is respected before work is done, not just when rows are
written: companies past the limit are no longer enriched only to be discarded.
- Named decision-makers from team, about and leadership pages, with job title and
LinkedIn profile, even when no email is published. Testimonials and reviews are excluded.
When the company publishes some personal addresses, their format is applied to other
named staff (
personEmails[].type: "pattern", with a confidence).
- The primary contact is always reachable: the best contact with an email address.
- Bot walls and parked domains are recognised. Cloudflare challenge and block pages are
no longer scraped into a company called "Attention Required!" with an IP address for a
phone number; parked, expired and for-sale domains are not billed.
- Company names are chosen by matching every candidate (schema.org, og:site_name, each
title segment, logo alt text, copyright holder) against the domain, including the domain
a site redirects to. Taglines and page names no longer become the company name, and
registered names ("… SRL", "… LLP") go to
companyLegalName.
- Locations: street addresses no longer land in the city field,
<br>-separated address
blocks are parsed, and the head office is preferred when several offices are listed.
- schema.org LocalBusiness subtypes (Dentist, AccountingService, Winery…) are read for
contact data and used to classify the industry;
email and telephone are used.
- Industry classification covers far more small-business verticals and matches whole
words (no more "cpa" inside "cpanel").
- Removed a fabricated headcount: a "100-500" range was inferred from the company's
email-security vendor and was wrong as often as right.
companySizeEmployeesRange now
only carries a range the company publishes.
- GitHub organisations are attached only when their listed website matches the domain.
- Social links are validated by hostname; dates and IP addresses are no longer phones;
questions and prose are no longer job titles; many more multilingual and functional
shared inboxes are recognised; duplicate contacts are merged; empty strings are
null.
Platform hardening. Measured on 30 companies: 30/30 succeeded, 0 failed, 156s, 0.043
compute units - roughly a third of the cost per result the Actor used to run at.
- Dataset writes are serialised. Called concurrently from every worker,
Actor.pushData stopped settling for the later companies. The request handler waited on
a write that never returned, was killed by Crawlee at its timeout, and a company that had
been enriched successfully was lost. Writes now queue one at a time, are time-boxed, and
are retried once against a fresh handle. A company whose row cannot be stored is reported
rather than dropped, and is never counted as billable.
- A company is bounded by one deadline shared across all providers, each receiving only
the time that remains. Previously each provider had its own timeout and the worst case
summed to more than the request handler allowed, so a slow company was killed and lost
everything it had gathered instead of finishing with partial data.
- Companies are never retried. Providers already retry their own requests; re-running a
whole enrichment costs twice and risked charging twice.
- Redundant DOM parsing removed. Each page was parsed about four times, including a
full serialise-and-reparse per call in
visibleText. Now one parse per page.
- Chromium is installed into a writable path in the image, so the browser fallback
actually works on the platform instead of silently degrading to HTTP-only.
- Status messages are throttled and never awaited, and named-dataset access is time-boxed
and self-disabling. Both could previously stall a company for minutes through client
retry backoff.
- Domains are checked for an A or MX record before any page is fetched, so a dead domain
no longer costs four browser page loads.
Fixes for three failures that only appeared once the Actor ran on the Apify platform.
- Chromium was missing from the image.
playwright resolved to 1.58.2, which looks
for a browser revision the base image does not ship, so every browser fallback failed
with "Executable doesn't exist at /pw-browsers/...". The browser is now installed
explicitly during the build. Before this fix the Actor silently degraded to HTTP-only
and returned no data for JavaScript-rendered sites.
- Writing to the auxiliary
errors dataset could stall a run. The call returned
"Insufficient permissions" and the Apify client retried it with exponential backoff for
roughly two minutes, pushing the request handler past its 120-second timeout. Crawlee
then retried the request and the company was enriched and pushed a second time, billing
one input row twice. Named-dataset access is now time-boxed, and the first failure
disables it for the rest of the run and falls back to logging.
- A retried request can no longer re-push a company that was already pushed, whatever
the cause of the retry.
Cost and data-quality rework, and a move to pay-per-result pricing.
- The Actor is now billed once per company enriched. One input row produces exactly
one dataset row regardless of how many contacts are found, so run cost is predictable
from the CSV alone.
- Renamed from "B2B Lead Enrichment (Free)" to "B2B Lead Enrichment".
- Output is now one row per company instead of one row per contact. Company fields are
no longer duplicated across rows.
person* columns are replaced by primaryContact* columns carrying the best contact
found. The full list moved to the contacts array with a contactCount alongside it.
techStack no longer contains SSL issuers, GitHub repository topics or source
languages. These moved to sslIssuer, githubTopics and programmingLanguages.
- The mail host moved out of
techStack into emailProvider.
- Pages are fetched over plain HTTP first; a browser is launched only for pages that
render entirely in JavaScript. Previously every page load ran through Chromium.
- Removed a fixed 1.5-second wait on every page load, replaced by waiting for content to
actually appear.
- Chromium is now launched lazily and shared, so runs that never meet a JavaScript-only
site never start a browser at all.
- Browsers block images, fonts, media and stylesheets — only the DOM is ever read.
- Default concurrency raised from 3 to 10.
- Provider time budgets are now enforced by the waterfall, so one hung request can no
longer hold a company open until the request handler times out.
Measured on a 12-company sample, per-company wall time fell from 2.55 s to 0.74 s and
browser page loads fell from 100% of fetches to roughly 10%.
- Sub-pages are discovered from the site's own navigation instead of guessed from a
fixed path list. The previous defaults never reached the contact page at all.
- Documentation placeholders are rejected. Addresses and phone numbers inside code
samples are excluded, along with known example identities. Previously a Stripe API
sample address and the placeholder number
1000000000 were published as real contacts.
- Shared inboxes are surfaced as labelled contacts and in
companyEmails, ranked by
sales relevance. They were previously collected and then discarded.
- Contacts now carry names, titles, seniority and department read from team and
contact page cards, or derived from unambiguous
first.last@ addresses.
- Locations are parsed rather than passed through. Office lists, countries in the city
field, combined "City, ST" values and non-places like "Remote" are handled correctly.
- Added
emailProvider, emailSecurity, emailMarketingTools, dnsProvider,
sslIssuer, sslValidFrom, sslValidTo, companyLogoUrl, companyGithubUrl,
companyEmails, companyPhone, programmingLanguages and githubTopics.
- Added
enrichmentScore and fieldsFilled for filtering thin records.
- Industry classification reads page headings as well as the description, weighted lower
so marketing copy cannot outvote it, and covers 18 non-technology sectors.
- Homepages are retried on the
www host and over HTTP before being given up on.
- schema.org parsing now walks
@graph structures and picks the richest organization node.
- DNS falls back to DNS-over-HTTPS when the system resolver is unreachable.
"San Francisco, CA" resolved to country Canada, because CA is both a US state code
and an ISO country code.
- A short page carrying schema.org or a meta description was treated as a JavaScript
shell and needlessly re-fetched in a browser.
contact-us@ style addresses were treated as individuals, and a nearby office heading
could be published as that person's name.
- Hyphenated surnames were dropped when deriving names from email addresses.
- "Chief Executive Officer" mapped to no department.
- Unit test suite covering email validity, contact extraction, name derivation, phone
validation, location normalisation, JavaScript-shell detection and sub-page discovery.