# Changelog of Google Maps Scraper + Email & Social Extractor (`yasaslive/google-maps-leads-email-enrichment`) Actor

- **URL**: https://apify.com/yasaslive/google-maps-leads-email-enrichment/changelog.md
- **Full Actor documentation**: https://apify.com/yasaslive/google-maps-leads-email-enrichment.md

## Changelog

All notable changes to this Actor are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the Actor uses [semantic versioning](https://semver.org/).

### \[1.0.0] - 2026-09-26

#### Added

- **Google Maps search with no browser.**
  - Reads the same internal `/search?tbm=map` JSON the Maps web app fetches (`)]}'`-prefixed positional arrays).
  - Paginates with `!8i<offset>` until `maxPlacesPerQuery` is reached or results run out.
  - Supports search queries and Maps search URLs (`startUrls`).
- **Output fields per place:**
  - IDs: `placeId`, `cid` (decimal string), `featureId`
  - Business: `name`, `category`, `categories[]`
  - Address: `address` plus flat `street`, `city`, `state`, `postalCode`, `countryCode`, `neighborhood`
  - Location: `lat`, `lng`, `plusCode`
  - Contact: `phone`, `phoneUnformatted` (E.164), `website` (tracking params stripped), `domain`
  - Reputation: `rating`, `reviewsCount`, `priceLevel`
  - Other: `openingHours[]`, `timezone`, `claimed`, `imageUrl` (1024×768), `googleMapsUrl`, `searchQuery`, `rank`, `scrapedAt`
- **Contact enrichment** with `CheerioCrawler`:
  - Crawls the homepage plus the best contact/about/team/impressum pages (multilingual), with fallback paths, up to `maxPagesPerWebsite`.
  - Emails come from `mailto:` links, visible text, JSON-LD, Cloudflare `data-cfemail` protection and `[at]`/`(dot)` obfuscation.
  - Phones come from `tel:` links and text, validated to E.164 with the place's country.
  - Social profiles: Facebook, Instagram, LinkedIn (company preferred), X/Twitter, YouTube, TikTok.
  - Junk emails are filtered (Sentry, Wix, placeholders, `@2x.png`, hashes, no-reply).
  - Email domains get an MX check with a per-run cache.
- **Email classification** as `genericEmails` / `personalEmails`, plus the `emailTypes: generic-only` option (GDPR).
- **`includeReviewsSummary`:** the 1–5★ distribution, taken from the place-details endpoint.
- **Pay-per-event pricing:**
  - Events: `actor-start` ($0.01), `place-scraped` ($0.0015), `place-enriched` ($0.005, only when a contact is found).
  - A budget guard with in-flight reservations stops cleanly and writes `budgetReached: true` to `STATS`.
- **Robustness:**
  - `STATS` record with categorized error counters (blocked / rateLimited / proxy / network / parse / other), logged every 30 s.
  - Soft-block detection: captcha, `/sorry/`, consent and Cloudflare challenges.
  - Session rotation on 401/403/429, and exponential backoff with jitter.
  - Place deduplication across queries.
  - State that resumes after a migration.
  - A raw payload saved to the KV store (`LAYOUT_ERROR_*`) when Google's layout changes.
- **SSRF guard** for business websites: private, loopback and link-local hosts are refused.
- **Tests:**
  - 188 vitest tests, including parser tests on 7 real Google payload fixtures and 3 contact-page fixtures.
  - A Google route-handler suite.
  - An enrichment integration test against a local HTTP server.
  - Budget tests that use an SDK-faithful fake charging manager.
- **Tooling:** CI (format, lint, typecheck, test, build, npm audit, gitleaks, Docker build), `SECURITY.md`, `RUNBOOK.md`, `docs/BLUEPRINT.md`.

#### Decisions and deviations (documented per the project conventions)

- **Node 22 instead of 20** (`apify/actor-node:22`). Node 20 reached end-of-life on 2026-04-30. Dev tooling (vitest 5) also requires ≥ 22.12.
- **Default Google proxy is `RESIDENTIAL`, not plain `{ useApifyProxy: true }`.** Google blocks datacenter IPs. The exit country is matched per query: inferred from the text ("Austin, TX" → US), or taken from `countryCode`.
- **Websites use a separate `enrichmentProxyConfiguration`** (datacenter by default). About 750 KB of HTML per enriched place over residential proxy (~$8/GB) would cost more than the $0.005 `place-enriched` fee. Decided with the owner.
- **Review counts need a cookie.** Google omits review counts from about 7 in 8 cookieless responses. Sessions persist cookies, and a response with no counts is retried up to 2 times on a warmed session.
- **Plus codes are computed locally** (Open Location Code, verified against Google-issued codes), because search results don't include them.
- **`maxResults` (default 50) is a total cap across all queries**, next to the requested `maxPlacesPerQuery` (default 50). The user conventions require `maxResults` with default 50; the input description tells users to raise it when adding queries.
- **Website retries are capped at 2**, and session rotations at 1. Crawlee otherwise retries connection errors up to 10 times regardless of `maxRequestRetries`, and one dead website held a whole run for about 2 minutes in testing.
- **Budget stops cancel only the Google crawler.** Website crawls always hold a budget reservation, so work that is already paid for finishes.
- **Individual reviews and reviewer names are never collected** (GDPR data minimization). The details fixture was stripped of review texts for the same reason.
- **`charge()` after `pushData()`**, as the conventions require, rather than `Actor.pushData(item, eventName)`. The Actor bills two events per item and needs per-item control over the enrichment event.

#### Known limitations

- Google Maps returns at most about 120 places per query; split big areas into several queries.
- Place URLs (`/maps/place/...`) aren't accepted as start URLs.
- Websites that render contacts only with JavaScript (SPA contact widgets) or behind forms yield no emails.
- `personal` email classification means "not a recognized role address". It isn't identity verification.
