# Changelog of US B2B Leads Scraper — Google Maps (`sleek_waveform/us-b2b-leads-google-maps`) Actor

- **URL**: https://apify.com/sleek_waveform/us-b2b-leads-google-maps/changelog.md
- **Full Actor documentation**: https://apify.com/sleek_waveform/us-b2b-leads-google-maps.md

## Changelog

### \[2.1] - 2026-07-12

#### Reliability

- **Maps navigation retry on Apify** — on `Page.goto` timeout, reinitialize browser and retry once (`core/navigation.py`).
- Fixes multi-query failures (e.g. nail salon query 2/3 on build 2.0.2).

#### Data quality

- **Phone from listing cards** — extract from phone button `aria-label`, `data-item-id tel:`, and `tel:` links in card JS.
- **Two-pass place URL phone enrichment on Apify** — collect all listing cards first, then visit place URLs for missing phones (salons hide phone until full place view).
- **`core/phone_extraction.py`** + tests.

#### Verified (build 2.1.6)

- QA: 3/3 leads, 3/3 phones (~14s).
- Salons: 42 leads, 3/3 queries (~258s); hair+nail 20 each; barbershop 2.
- Listing evaluate hang fix + phone enrichment cap when emails enabled.

### \[2.0] - 2026-07-12

#### Speed (Apify fast path)

- **Skip Playwright email fallback on Apify** — aiohttp only; avoids extra tabs + ~10–30s/query.
- **Tighter listing extraction** — 8s → 6s timeout per card on Apify (5s regressed on frozen DOM).
- **Circuit breaker** — stop after 3 consecutive listing timeouts when leads already collected.
- **Health check aborts frozen page** — don't burn full run budget timing out every card.
- **Fewer email HTTP paths + shorter aiohttp budget** on Apify.
- **Faster scroll** — shorter settle/pause; max 2 scrolls on Apify.
- **Cap listings processed** — at most `max_results + 3` cards per query.
- **Close stray browser tabs** between queries (fixes multi-query nav failures).
- **`core/performance.py`** + `tests/test_performance.py`.

#### Verified (salons cohort, build 2.0.2)

- 40 leads in ~140s (was 7 leads / 344s on 1.9.1; 0 leads / 547s on 2.0.1).

### \[1.9] - 2026-07-12

#### Fix — salons / use_proxy cohort (0 results)

- **Skip browser proxy for Google Maps on Apify** — datacenter proxy caused `Page.goto` 20s timeouts on every query; emails already use direct aiohttp.
- **Longer Maps navigation timeout on Apify** — 20s → 40s; listing wait 8s → 12s.
- **Local proxy fallback** — if browser proxy navigation fails, reinitialize without proxy and retry once.
- **`core/proxy_strategy.py`** + `tests/test_proxy_strategy.py`.

### \[1.8] - 2026-07-12

#### Performance & revenue (apify-improve)

- **Apify proxy for `use_proxy` on cloud** — free public proxies broke Google Maps (0 results on LA salons cohort). Now uses `Actor.create_proxy_configuration()` on Apify; free-list proxies remain local-only.
- **Deadline-aware multi-query scheduling** — `core/run_scheduling.py` skips remaining queries and Playwright email fallback when run budget is too low (Pattern C/E).
- **Regression tests** — `tests/test_run_scheduling.py`, `tests/conftest.py` for legacy `google_maps_scraper` imports.
- **Cohort inputs** — `.actor/test-salons-la.json`, `.actor/test-qa-prefill.json`.

### \[1.3] - 2026-03-12

#### Email Extraction Overhaul

- **Visible-text-only extraction** — strips `<script>`, `<style>`, HTML comments, and ALL HTML tags before regex matching. Eliminates false positives from JS code, CSS, CDN URLs, and HTML attributes
- **mailto: link extraction** — extracts emails from `href="mailto:..."` attributes (highest confidence source)
- **Domain-matching preference** — emails whose domain matches the website domain are preferred over random matches
- **Obfuscation detection** — catches `info [at] domain.com`, `info(at)domain.com`, `info AT domain.com` patterns
- **Junk domain blacklist** — filters CDN/tracking domains (parastorage.com, wixstatic.com, gstatic.com, cloudfront.net, etc.)
- **Stricter TLD validation** — TLD limited to 2-6 chars, domain name must be ≥2 chars (rejects JS like `n.d@a.length`)
- **More contact paths** — checks 8 paths: /contact, /contact-us, homepage, /about, /about-us, /team, /staff, /company
- **Playwright fallback** — for leads where aiohttp found no email, uses already-open browser to visit /contact with JS rendering (5s budget)
- **Increased aiohttp budget** — per-request timeout 5→8s, overall budget 8→15s, read limit 500KB→1MB

#### Test Results (local)

| Query | Leads | Time | RAM | Emails | False Positives |
|-------|-------|------|-----|--------|-----------------|
| plumbers in New York, NY | 10/10 | 29s | 256MB | 2/9\* | 0 ✅ |
| lawyers in Miami, FL | 10/10 | 29s | 356MB | 4/10 | 0 ✅ |

\*9 leads had websites; 2 had visible business emails

### \[1.2] - 2026-03-14

#### Data Completeness Overhaul

- **aiohttp email extraction** — replaced browser-based email scraping with concurrent aiohttp requests (zero browser memory, 2.4x faster)
  - Checks 5 paths per website: `/`, `/contact`, `/contact-us`, `/about`, `/about-us`
  - No proxy needed — direct HTTP with gzip encoding
  - \~75% email hit rate on reachable sites
- **Address enrichment** — parses city/state from search query (e.g. "plumbers in New York, NY") and appends to partial street addresses from listing cards
- **Word-boundary street suffix detection** — fixes false positives where business names containing "Plumbing", "Drive", "Court" etc. were misidentified as addresses
  - Uses `\b` regex boundaries instead of substring matching
  - Requires at least one digit in address text (street numbers)
- **Phone regex updated** — now matches `+1 XXX-XXX-XXXX` international format shown in Google Maps cards

#### Test Results (local)

| Query | Leads | Time | RAM | Phones | Addresses | Emails |
|-------|-------|------|-----|--------|-----------|--------|
| plumbers in New York, NY | 10/10 | 18s | 249MB | 10/10 | 10/10 | 3/10 |
| dentists in Austin, TX | 10/10 | 23s | 329MB | 0/10\* | 10/10 | 0/10\* |

\*Dentist listing cards don't show phone/website in Google Maps card view

#### Memory

- Default memory reduced from 1024MB back to **512MB**
- Peak RAM: 249-329MB — well within 512MB limit
- Version bumped to 1.2

### \[1.1] - 2026-03-12

#### Memory Optimization (fits in 512MB!)

- Reduced default memory from 2048MB to 512MB (min 512MB, max 2048MB)
- Peak RAM: ~245MB (full process tree incl. Chromium) — well under 512MB
- Removed `--single-process` Chromium flag (increases memory with Playwright)
- Added low-memory Chromium flags: `--disable-renderer-backgrounding`, `--disable-backgrounding-occluded-windows`, `--js-flags=--max-old-space-size=256`, `--disable-extensions`, `--disable-background-networking`
- Sequential email extraction with single reused page (vs. N concurrent pages)
- Updated pricing to $0.004/lead ($4/1000) — competitive with official Google Maps Scraper
- Actor version bumped to 1.1

#### Performance (target: <30s per query on Apify)

- **Single JS evaluate() per listing** — replaces 8+ sequential Playwright DOM queries, 10x faster extraction
- Replaced `fake-useragent` (network download at startup) with 5 hardcoded modern User-Agent strings
- Fast sponsored-listing detection via JS evaluate (was slow `inner_text()` with 2s timeout)
- Sequential email extraction with 12s budget, 3s per-lead timeout, reused page
- Reduced page load timeout from 60s to 20s
- Reduced per-listing extraction timeout from 15s to 8s
- Reduced email page timeout from 5s to 3s on Apify
- Reduced listing wait-for-selector timeout from 15s to 8s
- Reduced scroll-and-load sleep from 0.8s to 0.5s on Apify, max scrolls from 3 to 2
- Reduced consent dialog delays by 60-70%
- Increased `_human_delay` Apify scaling from 5x to 10x reduction
- Skip sponsored/ad listings that cause DOM hangs
- Added timing instrumentation to scrape() and apify_main()

#### Dependencies (image size reduction)

- Removed `pandas` + `openpyxl` (~80MB) — only used for optional local Excel export
- Removed `requests` — all HTTP uses Playwright or aiohttp
- Removed `fake-useragent` — replaced with hardcoded UA pool
- Removed `sqlite-utils` — unused (sqlite3 stdlib already used)
- Removed `email-validator` — unused (custom regex validation in validators.py)
- Total: 6 dependencies removed, ~100MB+ smaller image

#### Optimizations

- Skip SQLite/Exporter initialization in Apify mode (not needed)
- Removed `EMAIL_EXTRACTION_DELAY` in Apify mode (no per-site delay)
- Email extraction logs count of emails found

### \[1.0] - 2026-03-08

#### Added

- Google Maps scraping with Playwright stealth for US business leads
- Niche-pack labelling — tag every lead batch for easy filtering
- Business email extraction from websites (info@, contact@, sales@ only — personal emails filtered)
- US-only enforcement — non-US results discarded before output
- Rate limiting (5–10 second delays, max 6 requests/minute) built-in and cannot be disabled
- GDPR/EU detection — flags EU addresses and coordinates
- Multi-query support — run multiple search queries in one Actor run
- Pay-per-event billing via `Actor.charge("result-scraped", 1)`
- Output: name, address, phone, website, email, rating, reviews_count, category, GPS coordinates, niche_pack label
- Local CLI mode with CSV, JSON, and Excel export
- SQLite storage for local deduplication
- Proxy rotation support (optional, for large runs)
- 30+ compliance tests covering rate limiting, email filtering, EU detection, data validation
