# Changelog of Local Lead Finder Pro | $4/1K | Lead Score, Tech Stack, Pitch (`apivault_labs/local-business-lead-finder-pro`) Actor

- **URL**: https://apify.com/apivault_labs/local-business-lead-finder-pro/changelog.md
- **Full Actor documentation**: https://apify.com/apivault_labs/local-business-lead-finder-pro.md

## Changelog

### 1.1 — 2026-09-19

- Added auto, local-search, URL-enrichment, and combined workflows for Console and MCP.
- Added safe `maxResults=25` UI/MCP default plus sales, contacts, and full JSON presets while preserving legacy API output.
- Added six numbered form sections, typed Dataset fields, and official results/summary/errors output links.
- Empty or incomplete input now succeeds with `INVALID_INPUT` SUMMARY and ERRORS instead of failing the Actor.
- Corrected the deployment target so this project cannot overwrite Local Leads Basic.

All notable changes to this Actor will be documented here.

### \[2.3] — 2026-06-07

#### Added — contact-quality upgrade (closing the gap vs Apollo-style finders)

This release targets the one thing directory scrapers do badly: turning a
listing into a **reachable, personalised** contact. None of these features
copy the proprietary-database competitors — they make the most of public
website data, for free.

**🔗 Real social profiles** — `socialProfiles{}` now contains the *actual*
Facebook / Instagram / LinkedIn / Twitter-X / YouTube / TikTok / Yelp URLs the
business links to from its own site (not just Google *search* links). Share/
intent/dialog URLs are filtered out. The old `socialSearchUrls{}` stays as a
fallback for businesses with no site.

**👤 Owner / decision-maker name** — `ownerName` pulls the owner, founder or
principal from JSON-LD (`founder`/`author`) and visible patterns ("Owner: …",
"Founded by …", "Meet Dr. …"). The outreach pitch now opens with the owner's
**first name** when found ("Hi John —") instead of the business name — a
personalisation pure directory scrapers can't match. +5 lead-score bonus.

**✅ Free email deliverability scoring** — new `verifyEmails` (default `true`).
Every scraped email is scored 0-100 with **no paid API and no SMTP probing**
(which would burn sender reputation): syntax + **MX-record lookup over DNS-
over-HTTPS** + role-based detection (info@, sales@) + disposable-domain check +
business-domain match. Adds `emailDetails[]`, `bestEmail`, `bestEmailConfidence`
and `bestEmailStatus` (valid / accept-all / risky / disposable / invalid).
A deliverable email (≥70) gives +7 lead score and wins `bestContact`.

**📄 Multi-page contact crawl** — new `crawlSubpages` (default `true`).
**Cost-aware**: only when a homepage has no email does the actor also fetch
`/contact`, `/about` and `/team` (max 3 extra pages, stops at the first email).
`pagesCrawled` reports how many pages were read.

**🌐 Direct URL / domain mode** — new `startUrls` input. Paste your own list of
business website URLs (or bare domains) and the full enrichment pipeline runs
on them, skipping the YellowPages search entirely. Turns a raw CRM domain list
into a fully-enriched lead list (CSV-in → enriched-out). Works alongside or
instead of `category` + `location`.

**🏷️ Business description** — `businessDescription` grabs the company tagline
from meta description / og:description / JSON-LD.

#### Changed

- `category` + `location` are no longer strictly required — supply them, OR
  `startUrls`, OR both. The run fails only when neither is given.
- Lead score now rewards owner-name (+5) and deliverable-email (+7) signals.
- `bestContact` prefers the highest-confidence scored email.
- CSV export adds: `Owner Name`, `Email Confidence`, `Email Status`,
  `Business Description`, real `Facebook` / `Instagram` / `LinkedIn` /
  `Twitter/X` / `YouTube` / `TikTok` profile columns, `Pages Crawled`.
- Run summary adds `withDeliverableEmail`, `withOwnerName`, `withRealSocialProfile`.

#### New input parameters

- `startUrls` (array) — direct URL/domain enrichment mode
- `crawlSubpages` (default `true`) — crawl /contact + /about when homepage has no email
- `verifyEmails` (default `true`) — free MX-based deliverability scoring
- `extractOwnerName` (default `true`) — owner/decision-maker name extraction

#### Migration

Fully additive. v2.x output keys are unchanged; new fields appear alongside.
Disable any new behaviour via `crawlSubpages: false`, `verifyEmails: false`,
`extractOwnerName: false`.

### \[2.2] — 2026-05-22

#### Added — 5 more lead-quality features

**☁️ CloudFlare email decoder** — Many WordPress / CloudFlare-protected sites obfuscate emails inline as `data-cfemail="abc123..."`. The actor now decodes them with the standard XOR algorithm, recovering 20-30% more real emails on those sites.

**📱 Mobile-friendliness audit** — Detects 5 signals from website HTML head:

- `has-viewport-meta` (responsive design)
- `has-responsive-css` (`@media` queries)
- `has-mobile-alternate` (m. subdomain)
- `has-amp` (AMP version)
- `fixed-width-layout` (NEGATIVE signal)

Returns `mobileFriendly` boolean + `mobileSignals[]` per lead. Sites without mobile viewport get +8 to leadScore (clear pitch angle for "responsive redesign").

**🔍 SEO hygiene audit** — Checks 5 on-page basics:

- `hasMetaDescription`
- `hasOgImage` (Open Graph)
- `hasH1`
- `hasJsonLd` (structured data)
- `hasCanonical`

Returns `seoAudit: {hasX, ..., seoScore}` (0-100). Sites with `seoScore < 40` get +5 to leadScore (pitch SEO audit).

**🎯 Industry-specific outreach pitches** — 8 industries with custom angles:

- **Plumbers / Electricians / HVAC**: "emergency calls — half your jobs come from someone Googling at 2am"
- **Restaurants / Pizza**: "online menus and reservation links — diners decide where to eat from their phone"
- **Dentists**: "online booking — patients now expect to schedule a cleaning the same way they book an Uber"
- **Lawyers / Attorneys**: "Google rankings for '\[city] \[practice area] attorney' — that's where 80% of clients start"
- **Auto Repair**: "Google reviews + 'near me' — most car owners pick the nearest 4★ shop"
- **Salons / Hair / Barber**: "Instagram-style portfolio gallery — clients book based on photos"
- **Gyms / Fitness**: "membership signup forms + class schedules — gym shoppers compare 3-4 sites"
- **Landscaping / Roofing**: "before/after photo galleries — homeowners hire based on portfolio"
- **Real Estate / Realtor**: "IDX listings + lead capture — agents without sites lose 40% of online enquiries"
- **Cleaning Services**: "online quote forms + booking — most customers want a price in 60 seconds"

The industry angle is woven into the existing dynamic pitch (no-website / dead-site / Wix / SEO templates).

**📍 Geocoding via OpenStreetMap Nominatim** (opt-in)
New `enrichGeocode` parameter (default `false`). When enabled, every lead gets:

- `lat`, `lng` — decimal coordinates
- `geocodedAddress` — Nominatim's canonical formatting
- `osmUrl` — deep link to the location on OpenStreetMap

Useful for territory-routing in CRMs (Pipedrive territories, HubSpot deal regions) or plotting leads on a custom map dashboard. Off by default because Nominatim asks for ≤1 req/sec, so this slows runs by ~1s per lead.

#### Improved — email scraping accuracy

Cleaned up false positives that were leaking through earlier versions:

- TLD whitelist now rejects JS-fragment artefacts (`window.location.reload` no longer matches `loc@ion.reload`)
- Lookbehind `(?<![A-Za-z0-9.])` rejects emails extracted from URLs (`www.flavorplate.com` no longer matches `flavorpl@e.com`)
- Stricter `EMAIL_OBFUSC_RE` requires visible separators (brackets / parens / whitespace) so brand names like "flavorPLATE" don't match as "flavor\[PL]\[AT]\[E].com"
- Domain blacklist expanded with `parastorage.com`, `wixstatic.com`, `cloudfront.net`, `gravatar.com`, `wp.com`, `automattic.com`, etc.
- Local blacklist adds placeholder addresses (`user`, `youremail`, `yourname`, etc.)
- Pluggable `_is_plausible_email()` helper applied to every extraction path (mailto, CF-decoded, plain, obfuscated)

#### Added — CSV export columns

`Latitude`, `Longitude`, `Geocoded Address`, `OSM Map`, `Mobile Friendly`, `SEO Score`

### \[2.1] — 2026-05-22

#### Added — 5 more enrichment features

**📧 Real email scraping from website**
For every lead with an alive website, the actor now scrapes plain-text emails and `mailto:` links from the homepage HTML. Filters out 30+ tracker / CDN domains (Wix, Google, FB, Sentry), Retina image hashes (`image@2x.png`), and noreply addresses. Returns up to 5 unique emails per lead in `emailsFromWebsite[]`.

**📞 Real phone scraping + E.164 normalisation**

- New `phoneE164` field on every lead — listing phone normalised to `+1XXXXXXXXXX`
- New `phoneTel` field — `tel:+1XXXXXXXXXX` click-to-call URL ready for HTML or buttons
- New `phonesFromWebsite[]` field — additional numbers found in `tel:` links and page text on the website (different department, mobile, after-hours, etc.)

**🏷️ Chain / franchise detection**
\~50 national chain brands matched by name (Roto-Rooter, Subway, Domino's, Great Clips, Anytime Fitness, RE/MAX, Servpro, Verizon, AT\&T, etc.). Two new fields:

- `isChain` (boolean)
- `chainBrand` (string — e.g. `"roto-rooter"`)
  Chains get a -30 leadScore penalty (corporate marketing controls spend, hard to sell to). New input `excludeChains: true` drops them entirely.

**🎯 Best-contact-channel recommendation**
New `bestContact` field returns `{channel, value, label}` — picks the single highest-confidence outreach path so users don't have to scan 5 fields. Priority order:

1. Real email scraped from website
2. Email from YellowPages listing
3. Phone (E.164) from listing
4. Phone scraped from website
5. Website contact page URL
6. Email guess (verify before use)
7. Website homepage / Listing URL

**📅 Brand age via Wayback Machine**
Single tiny request to `archive.org/wayback/available` per lead returns the year of the first archived snapshot. New field `brandAgeYears`. Lead score bonus:

- `brandAge >= 5` AND `websiteAlive=false` → **+10** (established brand with dead site = prime replacement target)
- `brandAge >= 5` (alive) → +3

**🔗 Contact-page discovery**
New `contactPageUrl` field — auto-found `/contact` / `/reach-us` / `/get-in-touch` link on the website. Falls into `bestContact` priority when no email is found.

#### Changed

- Lead score now factors in chain status (-30) and brand age (+10/+3)
- CSV export adds 9 new columns: `Phone (E.164)`, `Phone Click`, `Best Contact`, `Best Contact Channel`, `Best Contact Label`, `Is Chain`, `Chain Brand`, `Brand Age (years)`, `Contact Page`, `Emails from Website`, `Phones from Website`
- Aggregate summary adds `chainCount` and `withRealEmailScraped` fields
- Email guesses now skip known directory aggregator domains so we don't generate `info@yellowpages.com` when only a listing URL is available

#### New input parameters

- `excludeChains` (default `false`) — drop chain franchises
- `enrichBrandAge` (default `true`) — Wayback Machine brand-age check

### \[2.0] — 2026-05-22

#### Added — major lead intelligence upgrade

Every record now arrives **sales-ready**, not just scraped.

**Lead scoring + tiers:**

- `leadScore` (0-100) — composite signal combining no-website / dead-site / DIY-builder / review count / rating / years in business / contact completeness
- `leadScoreReasons[]` — every contributing signal in plain English
- `leadTier` — `cold` (<35) / `warm` (35-54) / `hot` (55-74) / `on-fire` (75+)
- Results sorted by `leadScore` descending

**Website intelligence (`enrichWebsites=true` by default):**

- HEAD/GET probe of every lead's website — `websiteAlive`, `websiteStatus`, `websiteSslValid`
- Tech stack detection (`websiteTechStack[]`) for **12 platforms**:
  Wix, Squarespace, WordPress, Shopify, Webflow, GoDaddy,
  Weebly, WordPress.com, Joomla, Drupal, ClickFunnels, GoHighLevel
- Dead-site bonus: +25 to lead score (abandoned site = ripe for replacement)
- DIY-builder bonus: +15 (Wix/Weebly/GoDaddy/WordPress.com — replaceable)

**Outreach helpers:**

- `emailGuesses[]` — `info@`, `contact@`, `hello@`, `office@` from website domain
- `socialSearchUrls{}` — 1-click search links for Facebook, Instagram, LinkedIn, Google Maps, Google Search
- `outreachPitch` — auto-written 2-sentence cold opener, tailored to no-website / dead-site / Wix / Squarespace / generic scenarios. Uses business name, city, rating, review count when impressive

**Filtering & sorting:**

- `minLeadScore` input — drop cold leads at the source
- `maxResults` input — hard cap after sorting (cost control)

**Output formats:**

- `exportFormat: "default"` (full JSON, includes everything)
- `exportFormat: "csv"` — flat record with HubSpot / Pipedrive column names: `Company`, `Industry`, `Lead Score`, `Lead Tier`, `Outreach Pitch`, etc.
- `exportFormat: "both"` — full JSON plus a nested `_csv` field

**Aggregate summary:**

- Every run ends with one `_summary: true` record containing
  `totalLeads`, `withoutWebsite`, `withDeadWebsite`, `avgLeadScore`,
  `leadTierBreakdown`, `topTechStacks`, `category`, `location`

**Cost:** unchanged. All enrichment is included in the $4/1K Pro price.

#### Changed

- `categories` updated from `["LEAD_GENERATION", "JOBS"]` to
  `["LEAD_GENERATION", "BUSINESS", "MARKETING"]` (more relevant)
- Title now leads with the new value props: *"Lead Score, Tech Stack, Outreach Pitch"*
- Default lead is now the highest-scoring one — the actor sorts by `leadScore` descending

#### Migration

v1.x users keep working unchanged — the `Business Name`, `Phone`, `Address`,
`Rating`, `Reviews Count`, `Category`, `Website`, `Email`, `Hours`,
`Years in Business`, `Listing URL`, `hasWebsite` fields all remain.
The new fields are additive. Disable enrichment via:

```json
{
  "enrichWebsites": false,
  "enrichEmailGuesses": false,
  "enrichSocialUrls": false,
  "includeOutreachPitch": false
}
```

### \[1.0] — 2026-05-11

#### Added

- Initial release
- YellowPages scraper by category + location
- `hasWebsite` flag for "businesses without website" filtering
- `onlyWithoutWebsite` filter
- Multi-page parallel scraping
