# Changelog of Google Maps Email Extractor & Lead Scraper (`highbrow_fame/google-maps-email-extractor`) Actor

- **URL**: https://apify.com/highbrow\_fame/google-maps-email-extractor/changelog.md
- **Full Actor documentation**: https://apify.com/highbrow\_fame/google-maps-email-extractor.md

## Changelog

All notable changes to this Apify Actor. Built with transparent development — every version and its tradeoffs documented.

### Measured, not changed (2026-08-17) — image blocking, second and final attempt

No release. The live build is unchanged at 1.4.21; this records an experiment that was built,
measured and reverted, so nobody spends the money again.

**What prompted it.** A per-run cost breakdown showed residential proxy is **~69% of this
Actor's platform bill** — on a 50-lead run, $0.655 of $0.953. That corrects a long-standing
note claiming proxy was ~1%. Since the Actor reads text and never looks at a photo, blocking
images through an $8/GB proxy looked like free money.

**What was tried.** `--blink-settings=imagesEnabled=false` as a browser launch flag — the
browser-level approach the reverted v1.4.1 interception note explicitly recommended, with no
per-request Node round-trip. Built as tag `beta` (1.4.22).

**Result — no saving, and a new failure mode.** Same input throughout: "digital marketing
agency in Austin", `maxResults=25`, RESIDENTIAL proxy.

| | leads | proxy GB | total | $/lead | runtime |
|---|---|---|---|---|---|
| images on | 25 | 0.0284 | $0.312 | **$0.01247** | 348 s |
| images off #1 | 20 | 0.0228 | $0.251 | **$0.01253** | 283 s |
| images off #2 | 25 | 0.0780 | $0.733 | **$0.02930** | 443 s |

Run #1 looks like a 20% win until you count leads: it delivered 5 fewer. Its scroll loop gave
up early — `No new results after 5 scrolls. Total: 20` — where the baseline logged
`Collected 26 links (reached maxResults=25)`. Per lead it saved **nothing**. Run #2 did reach
25 leads, and burned 2.7× the baseline's bytes doing it, needing 33 links to get there.

**The finding that closes the question:** two identical image-free runs ranged 0.0228–0.0780 GB,
a **3.4× spread**. The run-to-run variance is several times larger than any effect being
looked for, so this is not a measurement that a few more runs would settle — and without
images the scroll heuristic stops being dependable, which costs delivered leads.

Reverted; tag `beta` re-pushed clean as 1.4.23. Total cost of the experiment: **$1.30**.

#### Follow-up: which phase actually spends the proxy bytes

Same day, same question from the other end. Two crawlers share one residential proxy —
Playwright on Google Maps, and CheerioCrawler on the businesses' own websites. Only the
second has an obvious cheaper option, since ordinary sites do not throttle datacenter IPs
the way Google does.

Attributed **inside single runs**, not by diffing them: the Cheerio side was tallied
in-process from response `content-length`, and the Playwright side obtained by subtracting
from Apify's authoritative per-run proxy total. In-run attribution was the point — with
3.4× variance between identical runs, a two-run diff could not have answered this.

| | total | Cheerio (websites) | Playwright (Maps) |
|---|---|---|---|
| run A, 25 leads, 373 s | 40.6 MB | ≤14.4 MB (36%) | 26.2 MB (**64%**) |
| run B, 20 leads, 479 s | 95.6 MB | ≤10.3 MB (11%) | 85.3 MB (**89%**) |

**Google Maps is the byte hog, and it is the unstable one.** Playwright ranged 26→85 MB
across two runs while Cheerio held steady at ~0.5 MB per lead. Run B is the instructive
one: fewer leads, more than double the bytes. Maps byte spend tracks how much a session
struggles and re-scrolls, not how many leads come out — which also explains the run-to-run
cost variance seen in the image-blocking test above.

The Cheerio numbers are **upper bounds**: 50 of 64 responses arrived chunked with no
`content-length`, so their decompressed size was counted instead. Real wire bytes are lower.

**Conclusion — the cheap win is smaller than it looked.** Moving only the website-fetch
phase to a datacenter proxy caps out at roughly a third of proxy bytes and realistically
less, while risking email recall on sites that already return 403 often enough through
residential IPs. Not taken. The money is in the Maps phase, where image blocking has now
failed twice; anything further there needs a design that stabilises the scroll loop first.

No release. Live build unchanged at 1.4.21; instrumentation removed, `beta` re-pushed clean
as 1.4.25. Cost of this measurement: **$1.06**.

### 1.4.21 (2026-08-13)

#### Fixed

- **An empty place render now gets one retry before it is discarded.** A navigation
  timeout always got a retry from Crawlee; a page that rendered empty (typically right
  after a consent wall on an EU exit IP) was discarded on the first attempt, silently
  costing the buyer a row they asked for. Measured on 2026-08-13: 2 of 5 places on a
  bad-exit-IP run rendered empty, while the run's one retried timeout request delivered
  fine — the retry path works, empty renders just never used it. Discarding (never
  billing) an empty row is unchanged; it now only happens after a fresh navigation
  attempt. Baseline for comparison: 0 of 5 discarded on 2026-08-10, 1 of 39 on 2026-08-11.

#### Added

- **Actor output schema** (`.actor/output_schema.json`): the Console Output tab now
  links the delivered leads as JSON and CSV instead of showing raw storage. One of the
  three improvements the Store quality profile (77/100 on 2026-08-13) explicitly lists.

### \[1.4.18] — 2026-08-06

#### Added — a closing line for sample-sized runs

The post-run review prompt is gated on `leads >= 10 && emailRate >= 0.3`. That bar is right:
a thin run does not deserve a star, and asking after one earns a bad review rather than no
review.

But the Store form **prefills `maxResults: 5`** (deliberately, so Apify's daily automated test
finishes inside its 5-minute window) while the `default` is 25. So a first-time user who opens
the listing and presses Start lands at 5 leads, below the gate, and the run closed with a
summary that never said this had been a sample rather than a market read.

Runs under 10 leads now close with one line saying exactly that, pointing at `maxResults`.
Runs with 10+ leads but a weak email rate stay silent on purpose: neither a review prompt nor
a "run bigger" nudge is honest when the output itself was thin.

No pricing figures in the message, so it needs no edit on 10 August.

### \[1.4.16 / 1.4.17] — audit fixes, 2026-08-05 and 2026-08-06

Found by a full two-repo audit. **Live since 2026-08-06** (1.4.17 tagged `latest`).

#### Fixed — runtime claims made consistent with the linked demo runs

The README carries three **publicly linked** demo datasets with stated run times (verified
2026-08-05, datasets still live at 25 / 20 / 20 items): Manhattan 25 = 4m15s, London 20 =
2m30s, Berlin 20 = 3m30s. Two runtime models were fitted to the only two hard measurements —
the Manhattan demo (25 leads = 255 s) and the CHANGELOG 168-lead run (1329 s):

```
runtime ≈ 70 s + 7.5 s/lead     (fits both points exactly)
```

This matches the code's own `perLeadSecs = 8` (4 maps + 3 email + 1 validate) minus the
`PREFLIGHT_SAFETY_FACTOR = 2` the estimator adds purely to provision the timeout — the
estimator is deliberately 2× conservative, so its "540 s / 9 min" for the default is a
timeout budget, not an expected runtime.

An earlier pass in this same session had pushed the small runs to ~6–8 min by reading the
estimator's padded output as the expected time. That **contradicted the linked demos**
(line 113 said ~8 min for the very run line 87 links at 4m15s). Corrected to the model:

| Location | now |
|---|---|
| `INPUT_SCHEMA.json:30`, `README.md:342` | default 25 → "about 4–5 min (longer when sites are slow), ~70s + 7.5s/lead" |
| `README.md:113` (Manhattan detail) | **4m 15s** — matches the linked demo exactly |
| `README.md:41–46` (cost table) | 25→~4, 20→~3, 100→~14, 500→~65 min |
| `README.md:506–514` (effective-rate table) | 20→~3, 25→~4, 100→~14, 200→~27, 300→~40, 1000→~2.1h |
| `INPUT_SCHEMA.json:30` trailing | "500 leads ≈ 40 min" → "≈ 65 min" (the 40 implied ~4.8s/lead, contradicting the same field) |

The 50-lead rows (~8 min) were kept: the model gives 7.4 min, so ~8 is a safe round-up.
The 168-lead measured point (22 min) sits between the 100 and 200 rows and validates the
whole curve.

#### Fixed — advice that contradicted the actor's own preflight

`INPUT_SCHEMA.json:30` recommended raising the timeout "to at least 7200s" for >200 results.
The live `defaultRunOptions.timeoutSecs` is **14400**, so the advice was half the default and
would have made things worse. It also advertised up to 1000 leads, which the preflight
refuses: at `2 × (70 + 8000) = 16140 s` the estimate exceeds the 4-hour default, and
`CHANGELOG.md:131` states the inverse — **4 hours tops out at 891 leads**. Now says exactly
that, and gives 18000 s as the figure needed for 1000.

#### Fixed — outbound marketing quoting a future price as current

`marketing/apify-featured-submission.md:71` and `marketing/KIKULDENDO-apify-featured.txt:41`
both said "Cost at current pricing: $1.24" for the 168-lead Austin demo. $1.24 is the
**10 August** rate; today it is $0.84. The `.txt` is the file that goes to Apify's featured
team and carried no caveat at all. Both now state both figures with dates.

ℹ️ `README.md:10` (hero line, "22 min run · $1.24 cost") was **deliberately not changed**.
It is wrong today and becomes correct on 10 August, and adding a change-then-revert step to
an unattended scheduled task is worse than five days of overstating our own cost.

### \[Measured, not changed] — maxConcurrency 5 → 20, 2026-08-03

`maxConcurrency` stays at its default of 5. Same paired method as the memory benchmark: both configurations started simultaneously, three pairs, identical input (`restaurants in Budapest`, `maxResults: 10`, residential proxy).

| Pair | conc=5 | conc=20 |
|---|---|---|
| 1 | 159.8 s / $0.2373 | 435.4 s / $0.4028 |
| 2 | 249.6 s / $0.3991 | 160.5 s / $0.2805 |
| 3 | 186.4 s / $0.3193 | 257.2 s / $0.4636 |
| **mean** | **198.6 s / $0.3186** | **284.4 s / $0.3823** |

Four times the concurrency came out 43% slower and 20% more expensive, losing in two pairs of three.

Output was identical in every pair: 10 leads on both arms, and the same email count (7/7/8). No blocking or data loss showed up, so the input schema's warning that lower values "reduce the chance of being blocked" was not what this measured. Higher concurrency simply does not buy throughput here.

**This is the direct check on the 1.4.12 preflight recalibration.** That change removed the `/ maxConcurrency` divisor on the grounds that the Maps feed scroll is serial. If the old formula had been right, conc=20 should have finished roughly 4× faster than conc=5. It finished 1.43× slower — the old assumption was wrong by a factor of about 5.7 on this axis alone.

Caveat worth keeping: run-to-run spread exceeds the effect. conc=5 ranged 159.8-249.6 s, conc=20 ranged 160.5-435.4 s. With n=3 this does not establish that 20 is reliably worse. It does establish that it is not faster, which is the part that matters.

### \[Measured, not changed] — memory 4096 → 8192, 2026-08-03

`memoryMbytes` stays at 4096. The 1.4.2 entry records that halving it to 2048 made the crawler 5.4× slower, because CPU on Apify scales with memory and CPU is the bottleneck. That left the obvious question open in the other direction, and a Creator-plan subscription the same day lifted the RAM ceiling from 16 GB to 64 GB, so it finally became testable.

Three paired runs, same input (`restaurants in Budapest`, `maxResults: 10`, residential proxy). Each pair started both configurations **simultaneously**, so time-of-day and proxy-pool conditions hit both equally — necessary, because the same input has measured anywhere from 124 s to 312 s today.

| Pair | Memory | Time | CU | Compute | Proxy | Total |
|---|---|---|---|---|---|---|
| 1 | 4096 | 170.4 s | 0.189 | $0.0379 | $0.3030 | $0.3457 |
| 1 | 8192 | 130.7 s | 0.290 | $0.0581 | $0.3199 | $0.3826 |
| 2 | 4096 | 163.6 s | 0.182 | $0.0364 | $0.2971 | $0.3382 |
| 2 | 8192 | **306.6 s** | 0.681 | $0.1362 | $0.4152 | $0.5575 |
| 3 | 4096 | 214.2 s | 0.238 | $0.0476 | $0.2587 | $0.3111 |
| 3 | 8192 | 135.4 s | 0.301 | $0.0602 | $0.2923 | $0.3572 |

8192 MB is genuinely faster: 23% in pair 1, 37% in pair 3. Pair 2's 306.6 s is a navigation stall of the kind seen elsewhere today, and it is worth noting that the extra memory did nothing to prevent it.

It still loses. Compute units are `GB × hours`, so doubling memory needs the run time to **halve** just to break even. Dropping the contaminated pair, 8192 MB is 31% faster and 38% more expensive in compute ($0.0592 against $0.0428). And speed earns nothing — users pay per delivered lead, not per second, so every millisecond saved is purely the developer's cost line.

Total cost moves less than compute alone, 11-15%, because residential proxy dominates the bill ($0.26-0.42 against $0.04-0.14) and is unaffected by memory.

Measured in both directions now: 2048 MB is much slower, 8192 MB is faster but not enough to pay for itself. 4096 MB is the local optimum. n=3 with one contaminated pair is thin, but the direction is not in doubt.

### \[Unreleased] — weekly health check, 2026-08-03

A local scheduled task, `apify-gmaps-health`, running Mondays 07:00 Europe/Budapest. It starts a real run with the exact Store prefill input and asserts five things, each mapping to something that actually broke on 2026-08-03: `SUCCEEDED` with exit 0; `runTimeSecs` under 240 s (against the platform's 300 s ceiling, so drift shows about a week before the badge would); a non-empty dataset with zero rows lacking name, phone and address, and at least one `primaryEmail`; a log free of `TypeError`, `ReferenceError`, `the reject button was not found` and `Still on Google's consent wall`; and a marketplace `notice` clear of maintenance, with any non-zero `TIMED-OUT` count reported.

`the reject button was not found` is the one that matters most — it is the tripwire for Google changing the consent button's `jsname`, which would otherwise turn every EU-exit-IP run into a silent empty result.

**Not the Actor Testing Actor**, which is what Apify's own docs recommend. [pocesar/actor-testing](https://apify.com/pocesar/actor-testing) is `FULL_PERMISSIONS` — it needs full account access to start runs on the owner's behalf — and the account owner declined that. The local task uses `curl` and the existing token, so it grants nothing new; the trade is that it only fires while the machine is on.

Weekly rather than daily because each execution starts a real run (~$0.21 measured) against a $5/month platform cap that customer runs also draw on. The task checks the remaining balance first and skips the run entirely below $0.50, rather than starving paying traffic.

### \[Unreleased] — unit tests, 2026-08-03

Not pushed; `tests/` is excluded via `.dockerignore` alongside `test-inputs/` and `scripts/`, so nothing new reaches the image. Ships with the next build.

`npm test` — **40 tests, ~0.2 s, no network, no platform credits.** Node's built-in runner (`node --test`), so there is no Jest and no new devDependency.

Coverage is deliberately narrow: the pure functions where a silent regression either costs money or corrupts every delivered record.

- **`cidFromPlaceId`** — the delta-mode dedup key. A real CID is `3899459889282892030`, well past `Number.MAX_SAFE_INTEGER`, so a Number-based implementation would round it and silently stop recognising duplicates — which means billing the customer a second time for leads they already own.
- **`normaliseWebsite`** — unwrapping Google's `/url?q=` click-tracking redirect. Miss it and the email scraper visits google.com instead of the business.
- **`parseRating` / `parseReviewCount`** — comma decimals (`4,5`) and space-grouped thousands (`12 603`). The multilingual crawl is the differentiator, so European formatting is the common case.
- **`rankEmails` / `pickPrimaryEmail` / `classifyEmail`** — which address becomes `primaryEmail`, including that a domain match must outrank a role prefix.
- **`gradeWebQuality`** — the modern/dated/poor thresholds a web agency filters on, and that a missing signals object stays `unknown` rather than collapsing into `poor`.
- **`computeLeadScore` / `leadReadiness`** — the advertised 0-100 score and its bucket boundaries.

#### The suite was mutation-tested, and that changed two of the tests

A suite that goes green on first run has not shown it can fail. Seven deliberate defects were injected into `src/utils.js` one at a time; **six were caught**, and the two misses were the useful part:

- `gradeWebQuality`'s modern threshold loosened from `>= 6` to `>= 5` passed everything. The boundary was asserted from one side only — there was a test for score 6, none for score 5. **A boundary asserted from one side is not pinned.**
- Raising the `Math.min(100, score)` cap to 200 changed nothing, because the weights sum to exactly 100 and the cap cannot fire. The test named "caps at 100" was verifying something unreachable. It now asserts that dropping a +5 signal from a maxed item moves the score by exactly 5 — which fails the day a new signal pushes the total past 100 and the cap starts silently swallowing the smallest ones. Verified against that exact injected change.

### \[1.4.14] — 2026-08-03

#### Changed — `maxResults` prefill 10 → 5, because one retry was enough to fail the daily test

Six platform runs of the Store prefill input, all 10 leads on residential proxy:

```
124.5   138.7   155   163   176.2   312.7   seconds
```

The automated daily test allows **300 s**. The last one blew through it, and the log says why:

```
WARN  Navigation timed out after 60 seconds.
      Reclaiming failed request back to the list or queue.
```

One place page hung, Crawlee retried it, and a ~150 s run became 312 s. Not a regression — the same run cleared 2 consent walls (identical to the 138.7 s run) and logged zero `h1` render failures, so the 1.4.10 and 1.4.12 changes are not implicated. It is ordinary network variance, and at 10 leads the margin was simply too thin to absorb it: ~150 s of work against a 300 s ceiling leaves no room for a 60 s stall plus a retry.

At 5 leads the expected run is ~110 s, which absorbs a stalled navigation and still lands inside the limit. `default` stays 25 — this only moves what the Store form pre-fills and what the daily test executes.

The trade is a thinner first impression for a new user, and it is worth taking: a low prefill costs one edit in the input form, while three consecutive failed daily tests earn an "Under maintenance" flag that removes the Actor from the marketplace.

### \[1.4.12] — 2026-08-03

#### Fixed — the preflight runtime estimate was 2.7-9.8× optimistic, so it let doomed runs start

The preflight exists to refuse a run that cannot finish inside its timeout — 0 events billed, clear remediation — instead of letting the user wait two hours for a partial dataset they still pay for. It could not do that job, because the estimate it refuses on was short on **every** run ever measured:

| Leads | Proxy | Actual | Old estimate | Ratio |
|---|---|---|---|---|
| 10 | residential | 124.5 s | 46 s | 2.7× |
| 10 | residential | 138.7 s | 46 s | 3.0× |
| 10 | residential | 155 s | 46 s | 3.4× |
| 10 | residential | 163 s | 46 s | 3.5× |
| 22 | residential | 456 s | 66 s | 6.9× |
| 168 | residential | 1329 s | 299 s | 4.4× |
| 10 | **datacenter** | 443 s | 45 s | **9.8×** |

Two structural errors. `estRuntime = 30 + leads × perLeadSecs / maxConcurrency` divided by concurrency as if the whole run parallelised — but the Maps feed scroll is serial, so the divisor invented speed that does not exist. And the 30 s fixed term was far below real startup; fitting the 10-lead and 168-lead points gives ~70 s fixed and ~7.5 s/lead marginal, the latter close to what `perLeadSecs` already computed.

The concurrency divisor is gone, the fixed term is 70 s, and two multipliers are applied:

- **`PREFLIGHT_SAFETY_FACTOR = 2`** — the 22-lead run landed 1.9× above the fitted line, so 2× is the smallest honest margin. The estimate is deliberately biased high: a refused run bills nothing and says what to change, a timed-out run bills the start fee and delivers a partial dataset.
- **`NON_RESIDENTIAL_PROXY_FACTOR = 2.5`** — identical input, 10 leads, 443 s on datacenter against a 124-163 s residential band. Google throttles datacenter exit IPs, and nothing in a lead-count model can see that.

The new estimate covers all seven measurements. What it changes in practice, against the 7200 s Actor default timeout: a 10-lead run estimates 5 min and a 25-lead run 9 min, so ordinary runs are unaffected. A 1000-lead run now estimates 4.5 h and is refused — it would previously have been waved through on a 27-minute estimate and then timed out at roughly 2.2 h.

**Why this surfaced now:** in the 30 days to 2026-08-03 the Actor logged **5 TIMED-OUT runs, against 0 that morning** — all after the `maxResults` default went 10 → 25 and doubled real runtime against an unchanged estimate. The v1.2.x note in this file records the same failure mode at 4 of 15 external runs; the preflight was written to stop it and never could.

#### Fixed — the refusal message recommended a `maxResults` that still would not fit

`maxResultsSafe` was left on the old arithmetic, so a refused run was answered with a cap derived from the same formula that had just been shown to be 3-7× wrong. It is now the exact inverse of the estimate: solving `available = 2 × proxyPenalty × (70 + tileOverhead + leads × perLeadSecs)` for `leads`. Verified at four timeout values — a 4 h limit yields 891 leads and back-computes to 14396 s, a 15 min limit yields 47 and back-computes to 892 s.

#### Fixed — the refusal advice was nonsense when the timeout was very short

Forcing the refusal path locally (`APIFY_TIMEOUT_AT` set to 200 s ahead) produced *"Lower maxResults from 100 to ≤ 1"* and *"Split into 100 smaller runs of ≤ 1 leads each"* — arithmetically correct, since a timeout that short does not cover startup at all, and completely useless as advice: 100 runs each paying full startup is strictly worse than one. When the budget does not cover startup plus one lead, options \[A] and \[C] are now suppressed and the message says plainly that lowering `maxResults` cannot help and the timeout has to go up.

### \[1.4.10] — 2026-08-03

#### Fixed — clearing the consent wall cost us the page load, and the empty row was billed

Found by running 1.4.9 on the platform rather than by reading it. That run cleared one consent wall and delivered 10 leads — **one of which was entirely null**: no name, address, phone, website or category, but a populated `googleMapsUrl`, `placeId` and `cid`, because those come from `request.url` rather than the page. It was a real restaurant, and the customer was billed for the empty row.

The log shows the mechanism in three lines:

```
17:45:27  Google consent wall detected — rejecting non-essential cookies.
17:45:30  Scraping place: .../Menza+Étterem+és+Kávéház/...
17:45:32  Extracted place: (no name) | website: none
```

Two seconds. `dismissConsent()` resolves its bounce-back on `domcontentloaded`, which consumes the page load Crawlee had already waited for — so the place panel had not rendered when extraction ran and every selector missed. The SEARCH handler never showed this because it waits on `div[role="feed"]` afterwards; the PLACE handler had nothing equivalent.

Two guards, because the first fixes this cause and the second covers the rest:

- After a wall is cleared on a place page, wait for `h1` before extracting.
- Never store a result with no name **and** no phone **and** no address. Whatever produced it — consent bounce, slow render, captcha, a layout change — it carries nothing the buyer can use, and `Dataset.pushData` is the `apify-default-dataset-item` billing trigger. Do not sell it.

### \[1.4.9] — 2026-08-03

Both fixes here are defects in 1.4.8, found by an adversarial audit of that build (31 claims raised across five independent review lenses, 13 surviving a refutation pass) and then reproduced locally before being touched.

#### Fixed — the consent fix only worked in English

1.4.8 matched the reject button by its label: `button[aria-label="Reject all"], button:has-text("Reject all")`. But the consent page renders in the `hl` language carried by the Maps URL, and `language` offers 40+ values. The 2026-08-03 verification happened to run `language: "en"`, which is the only reason it passed — every non-English EU run still failed exactly as before the fix, and the Store's daily test could never catch it because that runs `en` from a US exit IP where the wall never appears.

Reproduced before fixing — `language: "de"`, Hungarian IP, no proxy: *Consent wall present but no "Reject all" button found*, 0 places.

Probing `hl=de`, `hl=hu` and `hl=fr` showed the label changes but Google's internal `jsname` does not:

| `hl` | Visible label | `jsname` |
|---|---|---|
| de | Alle ablehnen | `tWT92d` |
| hu | Az összes elutasítása | `tWT92d` |
| fr | Tout refuser | `tWT92d` |

The locator now leads with `button[jsname="tWT92d"]` and keeps the English text only as a fallback for the day that identifier changes. Accept is `b3VHJd` — deliberately never matched. There is also **no positional fallback**: "the first of the two buttons" would eventually click Accept, and accepting on the user's behalf is not ours to do.

After: `de` 2 places, `hu` 2 places, `en` unchanged at 3 places / 3 emails.

#### Fixed — an uncleared wall billed the user for a junk row

The PLACE handler called `dismissConsent()` and ignored the result. On the two paths where the helper gives up — no reject button, or still on the wall after clicking — extraction continued against consent.google.com. Every business field came back null, but `googleMapsUrl`, `placeId` and `cid` are derived from `request.url` rather than the page, so the row looked legitimate on inspection. `Dataset.pushData` is the `apify-default-dataset-item` billing trigger, so the user was charged for it.

The handler now checks the URL after `dismissConsent()` and skips the request instead. Small money — $0.007 a row — but it also masked the empty-run diagnostic: enough junk rows and `placeResults.length` is no longer 0, so the "consent wall could not be cleared" status message never fires and the user gets a success CTA over a dataset of empty records.

#### Fixed — README understated cost on three demo lines

`$0.06` and `$0.07` for 20-lead runs, and `~$0.10` for a 25-lead run, against an actual `$0.20` and `$0.24` at the rates the same document quotes. A user reading `$0.06` and being billed `$0.20` is a support ticket at best.

### \[1.4.8] — 2026-08-03

#### Fixed — Google's consent wall silently produced empty runs

From an EU exit IP, Google answers a Maps URL with a redirect to `consent.google.com` ("Before you continue to Google Maps"). `div[role="feed"]` never renders, so the SEARCH handler waited out its 30-second timeout, logged a warning, and the run finished **`SUCCEEDED` with an empty dataset**. A user whose proxy did not yield a non-EU exit IP paid the run-start fee, received nothing, and got no error explaining why.

The Actor now detects the redirect and clicks **"Reject all"** — rejecting rather than accepting is deliberate: it is the privacy-preserving choice and reaches the feed just as reliably. Applied in both the SEARCH and PLACE handlers, because a fresh browser context carries no consent cookie and a place page can hit the wall even when the search that produced it did not.

Measured from a Hungarian residential IP, same input (3 results, no proxy):

| | Places collected | Emails |
|---|---|---|
| Before | **0** | 0 |
| After | **3** | 2 |

Cost on a non-EU exit IP is one URL comparison per page. Confirmed on the platform against the Store prefill input: `consentWallsSeen: 0`, run time 163 s versus 155 s for the previous build — within normal variance, no measurable overhead.

#### Changed — an empty run now says why, on the run itself

Zero results used to leave nothing but a line buried in the log. The run now carries a status message, which is what shows in Console, and it distinguishes the two cases that need completely different fixes: a consent wall that could not be cleared (change the proxy) versus a query that genuinely matched nothing (check the spelling).

Deliberately **not** `Actor.fail()`. A blocked consent wall is a proxy-configuration problem, and the Store's automated daily test treats a failed run as grounds for the "Under maintenance" flag. Same distinction the sibling AI-tracker Actor had to learn on 2026-08-02: configuration gap exits 0 with a message, genuine non-service fails.

#### Added — local test harness

Development tooling only; `test-inputs/` and `scripts/` are excluded via `.dockerignore`, so none of it ships in the image.

- `npm run check` — `node --check` plus JSON parse of the three schema files. Instant, free.
- `npm run test:local [name]` — copies `test-inputs/<name>.json` into the local key-value store and runs `apify run`. Inputs live outside `storage/` because the SDK wipes that directory on start. The runner refuses to launch if an input asks for the `RESIDENTIAL` proxy group. `smoke.json` exercises the full pipeline; `empty.json` covers the no-results path.

The consent fix is what made this useful: before it, every local run died at the search page with 0 places, so nothing downstream was reachable. It now runs the whole chain — place extraction, website crawl, email ranking, deliverability grading, web signals — from a laptop, for free.

**What it still cannot cover:** run time and cost. A laptop is not a 4 GB container, and the automated daily test allows 5 minutes with the `prefill` input. Only a platform run answers that, which is what `apify push --build-tag=beta` is for. Note also that `PROXY_EXTERNAL_ACCESS` is disabled on the free plan, so no Apify proxy group is usable from outside the platform.

### \[1.4.4] — 2026-08-03

#### Fixed — the Store's automated daily test would have failed

Apify runs every Store Actor daily with its **prefill** input and requires `SUCCEEDED` with a non-empty dataset **within 5 minutes**; three consecutive failures flag the Actor "Under maintenance", which removes it from Store visibility. This was measured, not assumed — a live run with the exact prefill input on build 1.4.3:

| | maxResults | Run time | Verdict |
|---|---|---|---|
| Before fix | 25 (schema default, no prefill) | **456 s = 7 m 36 s** | over the limit |
| After fix | 10 (`prefill`) | **155 s = 2 m 35 s** | passes with margin |

Cause: `maxResults` had `default: 25` and no `prefill`, so the automated test fell through to the default. The 10 → 25 default raise shipped in 1.4.2 for pricing reasons; nobody re-checked it against the 5-minute test budget. Live build until 2026-08-03 was 1.3.6 with a default of 10 (~3 min), so the regression arrived with 1.4.2 and would have landed on 10 August via the scheduled push either way — pushing 1.4.2 a week early is what surfaced it.

Fix is the one Apify's own docs prescribe: `"prefill": 10` on `maxResults`. `default` stays 25 for API callers that omit the field; only the Store form and the automated test see 10. Second run confirmed 10 leads, 7 with a primary email (70% hit rate).

#### Measured — residential proxy dominates cost on small runs, contradicting the 1.4.2 note

Full `usageUsd` breakdown from the two runs above:

| Line | 22-lead run | 10-lead run |
|---|---|---|
| `PROXY_RESIDENTIAL_TRANSFER_GBYTES` | $0.5374 (**83%**) | $0.3037 (**89%**) |
| `ACTOR_COMPUTE_UNITS` | $0.1014 (16%) | $0.0344 (10%) |
| everything else | $0.0091 | $0.0047 |
| **total** | **$0.6479** | **$0.3428** |

The 1.4.2 entry states proxy is "~1% of the bill" and uses that to argue against image blocking. **That figure read `DATA_TRANSFER_EXTERNAL_GBYTES`, which is not the residential-proxy line.** The actual proxy charge is `PROXY_RESIDENTIAL_TRANSFER_GBYTES`, billed at $8/GB on the free plan — 67 MB for 22 leads, ~3 MB per lead.

**Resolved the same day — this does not apply to user traffic.** Daily account usage for the 2026-07-13 → 2026-08-12 cycle shows `PROXY_RESIDENTIAL_TRANSFER_GBYTES` at **$0.0000 on every day except 2026-08-03**, where it is $0.8405 — i.e. the entire residential-proxy charge on the account is these two owner-run tests. Runs by other users put compute on the developer's bill but not residential proxy.

So the 1.4.2 conclusion stands for the economics that matter: compute dominates, proxy does not, and image blocking would not move the bill. Recorded here only so the next person who measures an owner-run test and sees proxy at 83-89% knows why it looks nothing like the Insights numbers.

### \[1.4.3] — 2026-08-03

#### Changed — the review prompt now waits for a good run

The end-of-run log printed `⭐ Found this useful? Rate the actor` on every run that delivered at least one lead, including thin ones. A prompt after a 3-lead run with no emails asks for a review at the exact moment the user has least reason to leave a good one — the downside is not a missing review, it is a bad one.

It now fires only when the run cleared **≥10 leads and a ≥30% primary-email rate**, and it prints as its own spaced block below the summary rather than as the third of three tip lines. Thresholds are deliberately conservative; the demo benchmark (168 leads, 53% email rate) clears both comfortably, a default 25-lead run on a well-covered market clears them, and a scraping run that went badly does not.

#### Changed — README review banner carries the user count

Replaced the generic *"if this actor saves you time"* opener with **"70+ teams have run this actor"** (measured 2026-08-03: 70 unique users, 395 total runs, 122 runs in the last 30 days at 96% success). Social proof is the constraint here, not the ask itself — the actor has 0 reviews and 0 bookmarks, and Store ranking runs on both.

#### Note — v1.4.2 shipped early, on purpose

v1.4.2 was built on 2026-07-27 but deliberately held back, because its README advertises the post-10-August rates. Holding it meant the price flip on 10 August depended on a scheduled local task running that afternoon; if the machine were off, users would have been charged $0.007 while the Store page still said $0.005.

It was pushed on 2026-08-03 instead, with a transitional note in the README and the `maxResults` input description stating that the live rate is $0.005 until 10 August. The Store card `description` and `seoDescription` were updated by API in the same pass to `$0.007/lead from 10 Aug` — true both before and after the boundary. **Verified after both pushes: `pricingInfos` is unchanged**, i.e. `apify push` does not touch pricing set in the Console. The transitional note is the only thing left to remove on 10 August.

### \[1.4.2] — 2026-07-27

#### Fixed — the Actor was losing money on small runs

Apify's Insights → Monetization for July 2026 showed, **for this Actor alone**, cost $0.93 against **profit −$0.42** — i.e. revenue of $0.51 covering 55% of what the Actor cost to run — plus a reported average **cost of $14.37 per 1,000 results**. (Account-wide the same month was revenue $0.90 / cost $0.94; the $0.90 includes the sibling AI-tracker Actor and must not be compared against this Actor's cost line.) Diagnosis from three real runs via `GET /v2/actor-runs/{runId}`:

| Cost line | Share of bill |
|---|---|
| `ACTOR_COMPUTE_UNITS` | **89–97%** |
| `REQUEST_QUEUE_WRITES` | 1–8% |
| `DATA_TRANSFER_EXTERNAL_GBYTES` (proxy) | ~1% |

Proxy traffic was never the problem, so no proxy change was warranted. More importantly, **healthy runs were never the problem either** — measured cost per 1,000 delivered results on real runs was **$1.97, $2.25 and $4.16**, i.e. $0.002–$0.004 per lead against a $0.005 price. The $14.37 average is an artefact of *small* runs: browser launch and Maps session warm-up cost the same whether a run returns 10 leads or 500, and that fixed cost was being recovered through a purely per-lead price. The `maxResults` default of 10 made the most common first run structurally unprofitable no matter what the per-lead rate was.

Raising the per-lead price could not fix this. The Store market for email-enriched Google Maps leads runs **$0.002–$0.010** (scraper-mind $0.004, faisalrjbd $0.005, herus13 $0.007, slothtechlabs $0.008); pricing above that band would simply end the runs.

#### Changed — pricing is now a two-part tariff

- `apify-default-dataset-item` **$0.005 → $0.007** — mid-band, still under the premium tier despite shipping MX/SPF/DMARC validation, lead scoring, multilingual contact-page crawl and delta mode.
- `apify-actor-start` **$0.00005 → $0.015** per GB of run memory ($0.06 on the 4 GB default) — sized to the measured fixed cost of a run (~$0.02–0.04 of browser startup). Amortises away with scale: effective rate is $0.0100/lead at 20 leads, $0.0071/lead at 1,000.
- `maxResults` default **10 → 25** — a 10-lead default charged $0.05 against more than that in fixed overhead.

Against the 168-lead Austin benchmark the new pricing yields $1.24 revenue. Margin depends on which measured cost rate you assume, and the three runs varied by 2.1× ($1.97 / $2.25 / $4.16 per 1,000), so quoting a single figure would be misleading:

| Cost rate | Platform cost on 168 leads | Margin |
|---|---|---|
| Best observed ($1.97/1,000) | $0.331 | 73% |
| Mean of three ($2.79/1,000) | $0.469 | **62%** |
| Worst observed ($4.16/1,000) | $0.699 | 43% |

**Plan against ~62%, not 73%.** The spread itself is the point: per-run cost is volatile, so the two-part tariff has to survive the worst case, not the best one. It does — 43% is still a working margin, which the old $0.005 flat rate was not.

#### Changed — compute (the safe subset)

- `maxRequestRetries: 1` on the **website** and **Facebook-enrichment** Cheerio crawlers. Dead, parked and blocked domains do not become reachable on the third attempt, and they have no email to find either way, so the hit rate is unaffected. The Maps crawler keeps Crawlee's default retries — its requests carry the actual yield.
- 2 MB ceiling on HTML passed to the email/social regexes. Oversized documents are truncated, not dropped, so a bloated page still yields whatever sits near the top.

#### Reverted — a compute experiment that measurably backfired

Three changes were shipped to build `1.4.1` on the `beta` tag and benchmarked against the 168-lead Austin run, then reverted. Recorded here so nobody tries them again:

| | Baseline v1.3 (4 GB) | Beta v1.4.1 (2 GB) |
|---|---|---|
| Playwright avg ms / request | **5,605** | **30,466** (5.4× slower) |
| Navigation timeouts | **0** | **8** |
| Places extracted in ~19 min | 168 (in 16 min) | 55 |
| Memory-overload samples | 0 | 0 |

- **Aborting `image`/`media`/`font` in a `preNavigationHooks` route** — the dominant cause. A catch-all `page.route('**/*')` round-trips every one of the hundreds of requests a Maps page issues through Node; that overhead dwarfs the bytes saved, and proxy bytes are ~1% of the bill to begin with. If this is ever revisited, block at the browser level with `--blink-settings=imagesEnabled=false`, which costs no per-request interception.
- **`memoryMbytes` 4096 → 2048** — the allocation looked wasteful (peak usage 1,224–1,482 MB, 32–36%), but on Apify CPU scales with memory, and the autoscaler logs show **CPU** overloaded and **memory never** overloaded in either configuration. Halving memory halved the CPU share; the throughput loss more than cancelled the halved CU rate.
- **`maxOpenPagesPerBrowser` 1 → 4** — no measurable benefit once the above dominated; reverted rather than left in unverified.

Net: the compute path is byte-for-byte v1.3 apart from the two retry/ceiling changes above. The economics fix is the pricing structure, not the crawler.

#### Note

Pricing increases carry Apify's mandatory 14-day notice to existing users; the new rates take effect after that window.

### \[1.3.5] — 2026-04-29

#### Added — Store SEO + comparison surface

- README **FAQ** section: 9 evergreen long-tail questions answered (differentiation vs other Google Maps scrapers, pricing model, email hit rate per market, multilingual coverage, preflight check, CRM/AI usage, recurring scrape patterns, what the actor does NOT do)
- README **competitive comparison table**: feature-by-feature matrix vs Compass GMS, Lukas Krivka, and others — explicitly identifies our differentiators (MX/SPF/DMARC validation inline, multilingual contact-page crawl, lead scoring 0-100, delta mode $0 dupes, preflight budget check, AI-ready outreach profile, Meta Ad Library URL)
- README **"Use the right tool for the job"** matrix — honest steering: Compass for million-record raw scrapes, this actor for validated EU multilingual outreach pipelines

#### Changed — Store metadata SEO hygiene

- `seoTitle`: "Google Maps Email Extractor with Built-in Email Validation" (58 chars, fits Google 60-char display)
- `seoDescription`: keyword-dense 143-char description with MX/SPF/DMARC + lead scoring + multilingual crawl + $0.005/lead + $0 duplicates
- `description` (Store card): 277 chars covering the full differentiator stack with localized contact-page keywords (`kapcsolat`, `kontakt`, `contacto`, `contatti`)

#### Note — PPE pricing went live

PAY\_PER\_EVENT pricing activated 2026-04-28T11:14 UTC — flat $0.005 per delivered lead + $0.00005 per run start. Failed/timed-out runs now cost $0. No code change in this release, just confirmation.

### \[1.3.4] — 2026-04-23

#### Added — Store listing visuals (README inline images)

- Three screenshots embedded in the README, hosted on the Apify CDN via a public key-value store (`lvBwYNZ1MRj8eWdsg`):
  - **Hero dataset table** — 168 Austin TX dentists, deliverability-graded, lead-scored, with readiness chips
  - **Preflight log** — actor run showing `[estimate]` + `[preflight]` block before any events are billed
  - **Record detail** — JSON record view of a single high-score lead with `emailValidation`, `webSignals`, `suggestedOpener`, `metaAdLibraryUrl`, `cid`
- Live Austin demo dataset referenced in README and marketing docs: `wMnqRj2ChH4NbsuVk` (168 records, 53% email hit, 25% high-deliverability, 83% hot leads, $0.84 cost, 22 min runtime)
- Store pictureUrl set to branded actor icon (GM monogram + VALIDATED checkmark, 512×512) — uploaded via programmatic Console automation; served from Apify images CDN

#### Fixed — Store metadata hygiene

- Cleared `UNDER_MAINTENANCE` notice flag (was blocking the Store banner)
- Replaced placeholder `exampleRunInput` (`{helloWorld: 123}`) with a real `dentist Austin Texas` 9-field config
- Categories set to `LEAD_GENERATION + ECOMMERCE + MARKETING` (3 is the API hard cap)

### \[1.3.2] — 2026-04-23

#### Changed — pre-launch hardening (before PAY\_PER\_EVENT goes live 2026-04-28)

##### Preflight runtime-estimator now tile-aware

Before this build, the preflight estimate ignored `geoGridTiles` multiplier — a user setting `geoGridTiles=5, maxResults=100` saw "~3 min" when the actual runtime was ~56 min. On the default 2h timeout this was harmless, but on custom short timeouts it could false-green an impossible run.

**New formulas** (mirroring the v1.3.1 verified Budapest test case within 1% accuracy):

```
tileUniqueFactor    = N > 1 ? N² × 0.75 : 1      // 75% unique survival after cross-tile dedup
estTotalLeads       = maxResults × (Q × tileUniqueFactor + nonTiledStartUrls)
tilePhaseOverhead   = N > 1 ? N² × 60s / min(concurrency, N²) : 0
estRuntimeSecs      = 30 + tilePhaseOverhead + estTotalLeads × perLeadSecs / concurrency
```

The 75% survival rate is calibrated from the v1.3.1 end-to-end test: 3×3 Budapest tiles discovered 90 raw hits, cross-tile dedup filtered 23 (25.5%), delivered 67 unique. New formula predicts 68 for that input — **off by one**, well within estimator tolerance.

**User-visible effects:**

- Tile-enabled runs log a dedicated line: `[estimate] geoGridTiles=5 → 1×25=25 tile searches, ~1875 expected unique lead(s), runtime ≈ 56 min (+5 min tile-scroll overhead).`
- Preflight refusal error message now enumerates `geoGridTiles` as a reduce-knob option `[T]` alongside existing `[A]..[D]` (lower maxResults / raise timeout / split runs / disable enrichments).
- `maxResultsSafe` calculation in the error message accounts for tile fan-out — it now recommends a value that actually fits.

##### Delta mode: CID as secondary dedup key

Delta mode (`sinceDatasetId`) previously dedup'd purely on `placeId` (format `0xHEX:0xHEX`). This worked, but had a failure mode we wanted to close before scaling: if Google Maps changes the `placeId` token format between runs (observed once on the v1.1.x → v1.2.x transition — capitalisation drift), the stable identifier is actually `cid`.

**New behaviour:** Delta mode loads BOTH `placeId` and `cid` from the prior dataset and skips a place if EITHER matches. CID is extracted at enqueue time from the place URL's `!1s0xHEX:0xHEX` token via `extractCid()` + `cidFromPlaceId()` (same helpers introduced in 1.3.1), so the check happens before any detail-page fetch. Zero extra cost.

Log line updated: `[delta] Loaded N placeId(s) + M CID(s) from prior run — will skip duplicates by either key.`

##### autoExtend × tiles — log-message clarity

The auto-extend path now explicitly calls out `geoGridTiles` in the warning so users can see why the estimated runtime is what it is:

```
⚠  Your run's timeout (60 min) is too short for maxResults=100 + geoGridTiles=5 (25 tile searches) (~56 min needed).
   autoExtend=true → starting a fresh run with 7200s (2h) timeout so you don't have to configure anything.
```

No functional change — the auto-extend logic already spread the entire input (including `geoGridTiles`) into the spawned run and used the tile-aware `estRuntimeSecs` for the buffer calculation. Code review confirmed correctness; this change is purely for observability.

#### Fixed

- `src/main.js`: preflight formula now includes `tileUniqueFactor` and `tilePhaseOverheadSecs` when `geoGridTiles > 1`.
- `src/main.js`: `maxResultsSafe` divisor corrected for tile fan-out.
- `src/main.js`: delta-mode bootstrap loads `cid` from prior dataset alongside `placeId`.
- `src/routes.js`: new `setSkipCids()` export; SEARCH handler dedup loop checks both placeId and CID.

#### Not changed

- PAY\_PER\_EVENT schema — still `apify-default-dataset-item: $0.005` + `apify-actor-start: $0.00005`.
- Tile logic, CID extraction helpers — stable from 1.3.1.
- User-facing API — `sinceDatasetId` usage is unchanged; the CID check is purely additive.

#### Validation

- Preflight math unit-tested across 6 scenarios (baseline, Budapest 3×3 verified, Manhattan 5×5, NYC 10×10, multi-query, unlimited).
- 3×3 Budapest prediction: 68 leads → actual in 1.3.1 test: 67 leads. **0.7% prediction accuracy.**
- `node --check` passes on all modified files.

### \[1.3.1] — 2026-04-23

#### Added — Tier 2 power features: geo-grid tiling + CID + cross-query dedup

##### 🎯 Geo-grid tiling — bypass Google's 120-result cap

Google Maps returns **at most ~120 places per search**, regardless of how many actually match. For small cities this is fine; for Manhattan (800+ restaurants), London (1200+ dentists), or NYC-scale coverage, you never see the long tail.

The new **`geoGridTiles`** input (1–10, default 1) splits the search area into an N×N geographic grid, issues the same keyword query against each tile's map viewport, and de-duplicates overlapping results by `placeId` so the final dataset is a single clean list.

**Recommended grid sizes:**

- `1` (default) → single viewport, ~120 max. Good for small towns + quick runs.
- `3` → 9 tiles, up to ~600 unique leads. Mid-size cities (Budapest, Lyon, Porto).
- `5` → 25 tiles, up to ~1500 unique leads. Big cities (Berlin, London, Chicago).
- `10` → 100 tiles, up to ~6000 unique leads. Mega-cities (NYC, Tokyo, Seoul).

**How geocoding works:** The query parser detects `"{term} in {location}"` / `"{term} near {location}"` / `"{term} {location}"` patterns, then calls free OpenStreetMap Nominatim to get the location's bounding box. No API key, no rate-limit budget to manage (we issue 1 geocode per query). If the location isn't parseable or Nominatim can't find it, the run falls back to an untiled single search — no silent failures.

**How viewport-zoom is chosen:** Each tile's edge length (in degrees) is mapped to a Google Maps zoom level via the empirically-calibrated formula `zoom ≈ log2(0.7 / latSpanDeg) + 10`, clamped to \[3, 19]. This reproduces published benchmarks (zoom 13 ≈ 15 km edge, zoom 14 ≈ 7 km, zoom 15 ≈ 3 km, zoom 16 ≈ 1.5 km) within ±0.5 zoom.

**Billing is linear in unique leads, not in tiles.** 5×5 tiles discovering 1500 unique places costs $7.50 + $0.00005 — same per-lead rate as a 1-tile run. You pay for data, not for the crawl budget.

##### 🔑 Google CID extraction + cross-tile dedup

New `cid` field on every dataset record — the **stable numeric Customer ID** Google uses internally for a business. Extracted via two paths:

1. `?cid=<DECIMAL>` — share URLs and some redirects already expose it.
2. `!1s0xHEX:0xHEX` feature-ID token — second hex group converted to decimal via BigInt (64-bit safe).

**Why it matters:**

- **Cross-query dedup**: The same business discovered via two different search terms (`"pizza in Manhattan"` + `"italian restaurants Manhattan"`) always has the same `cid`. Use it as your primary key when merging multiple runs.
- **Share URL reconstruction**: `https://maps.google.com/?cid={cid}` opens the place page regardless of language/region — the `placeId` hex-pair format sometimes drifts between Google runtime updates.
- **Tile dedup insurance**: When `geoGridTiles > 1`, neighboring tiles frequently overlap. The SEARCH handler now maintains a run-scoped `enqueuedPlaceIds` Set that short-circuits the second enqueue of the same place — reported as `[tile-dedup] Skipped N place(s) already enqueued from other tiles` in the run log.

##### 🧰 Internals

- **`src/tiles.js`** (new, ~200 lines): `parseSearchQuery`, `geocodeLocation` (Nominatim), `generateTiles`, `pickZoomForTile`, `buildTileSearchUrl`, `expandQueryToTiles` convenience wrapper.
- **`src/utils.js`**: `extractCid(url)` and `cidFromPlaceId(placeId)` helpers; `buildResultItem` now populates the `cid` field automatically (falls back to deriving from placeId if caller didn't supply it explicitly).
- **`src/routes.js`**: CID extracted in LABELS.PLACE handler from both `request.url` and `page.url()` (Google redirects /maps/search/ → /maps/place/ mid-navigation). New `getTileDedupCount()` export for the post-run summary.
- **`src/main.js`**: `geoGridTiles` destructured from input, SEARCH initial requests loop now uses `expandQueryToTiles`. Tile-aware `uniqueKey` protects against Crawlee collapsing distinct tile viewports. Post-run summary reports `tileDeduped` and `withCid` counts.

#### Added — dataset schema

- `cid` field in `storages.dataset.fields` of `.actor/actor.json`.

#### Added — input schema

- `geoGridTiles` integer field in `INPUT_SCHEMA.json` (1–10, default 1) with a detailed description including grid-size recommendations and the billing-is-linear-in-leads note.

#### Not changed (intentionally)

- PAY\_PER\_EVENT schema — still `apify-default-dataset-item` $0.005 + `apify-actor-start` $0.00005. Tile runs don't introduce new events.
- Delta mode (`sinceDatasetId`) — still dedups by `placeId`; cross-run `cid`-based dedup is deferred to a later version (would require backfilling CIDs in legacy datasets).
- Preflight + auto-extend — runtime estimator doesn't yet account for tile multiplication; users setting `geoGridTiles=5, maxResults=100` should expect up to 25× the single-tile runtime. Estimator update tracked for v1.3.1.

#### Migration notes

- Existing runs are unaffected (geoGridTiles defaults to 1 = untiled behaviour).
- Consumers relying on `placeId` as a primary key can continue to do so. `cid` is an additive field.

### \[1.2.15] — 2026-04-23

#### Changed — pricing schema aligned with Apify Console submission

The user submitted the PAY\_PER\_EVENT pricing through the Apify Console using the standard events, not the custom events defined in `.actor/actor.json`. Effective date: **2026-04-28 11:14 UTC** (5-day notice period — standard for Free → PPE transitions, not the 14-day "major change" window).

**Submitted events (going live Apr 28):**

- `apify-default-dataset-item` → **$0.005** per dataset record
- `apify-actor-start` → **$0.00005** per run start

**My old custom events in actor.json (never active, since Apify bills on the Console-submitted schema only):**

- `place-scraped` → $0.003
- `website-scraped` → $0.002

#### Why this matters

Once the new pricing goes live on Apr 28, any `Actor.charge({ eventName: 'place-scraped' })` or `Actor.charge({ eventName: 'website-scraped' })` call would attempt to charge an event that isn't in the pricing schema — best case: no-op (free runs forever, revenue = $0), worst case: SDK throws and crashes the run. Either way broken.

#### Fixed

- **`src/main.js` — 3× `Actor.charge()` custom-event calls removed:**
  - Line 314: `place-scraped` (in `requestHandler` after PLACE requests)
  - Line 445: `website-scraped` (after scraping each business website)
  - Line 574: `website-scraped` (after Meta Ad Library Facebook page fetch)
- **Replacement mechanism**: The platform now auto-fires `apify-default-dataset-item` for every record written by `Dataset.pushData` (batch loop at `main.js:654`). One push call with a 100-item batch = 100 billing events = $0.50. `apify-actor-start` fires once per run start automatically. **Zero explicit `Actor.charge` calls needed.**

#### Updated — `.actor/actor.json` pricingPerEvent

Replaced the custom-event schema with the Console-submitted one. `eventDescription` now documents the flat per-lead all-inclusive pricing, reinforcing the "no surprise bills, no per-subpage fees" message.

#### Updated — user-facing copy

- **`INPUT_SCHEMA.json`** `maxResults.description`: added explicit billing example (`100 leads = ~$0.50, 1000 leads = ~$5`).
- **`README.md`** — `💸 Transparent pricing` callout rewritten for the $0.005 flat model; above-the-fold cost table updated (`$0.10 → $0.125`, `$0.06 → $0.10`, `$0.07 → $0.10` for the small runs where the 2× rounding effect is visible). `How much will it cost to scrape {city}?` section rewritten with new formula `(leads × $0.005) + $0.00005`. Cost architecture note now emphasizes the flat-rate no-surprise-bills guarantee even when we crawl multiple contact subpages.

#### Not changed (intentionally)

- Preflight refusal logic (v1.2.9) — still fires before any `Dataset.pushData` call, so on refusal the dataset stays empty and 0 events fire. "0 events billed" guarantee preserved.
- Auto-extend logic (v1.2.12) — spawned runs still fire their own `apify-actor-start` ($0.00005 × 2 = $0.0001 total overhead if auto-extend kicks in), leads delivered on the spawned run fire `apify-default-dataset-item` normally.
- Delta mode (`sinceDatasetId`) — duplicate filtering happens **before** `pushData`, so no events fire for duplicates = $0 for already-known leads. Behaves exactly as before.

#### User-facing impact

Effective price is the same as before in the typical case (email scraping on): **$0.005 per lead**. Users who disabled `scrapeEmails` previously paid $0.003; now they pay $0.005 regardless (simpler mental model, still cheaper than any competitor selling per-email validation separately).

### \[1.2.14] — 2026-04-23

#### Added — Tier 1 market-leader setup (from 4-agent research synthesis)

- **SEO-optimized actor title**: `Google Maps Email Extractor & Lead Scraper` (42 chars — fits SERP / Apify Store tile). Down from the 73-char marketing subtitle that Google truncated mid-word.
- **Sub-300-char description** (API-enforced limit): hits the pain points Apify users search for — validated emails, delta mode, lead scoring, AI-ready profiles, PAY\_PER\_EVENT no-charge-on-failure.
- **5 dataset views in `.actor/actor.json`** — filtered presets the user can switch between in the Apify Console dataset UI, without writing any code:
  - `overview` — core business details (name, address, phone, email, rating)
  - `leadGeneration` — outreach-ready fields (validated email, score, AI profile, suggested opener)
  - `webAgency` — web-quality / tech-stack targeting for redesign pitches
  - `socialAds` — Facebook page ID + Meta Ad Library URL for active-advertiser filtering
  - `restaurantsFocus` — cuisine, price level, booking, menu, service options
- **README restructure for Apify Store SEO + first-click conversion**:
  - Status badge at top (`apify.com/actor-badge`) — visually signals "living, maintained actor"
  - **Above-the-fold 5 unfair-advantage bullets**: email deliverability grading, delta mode, Meta Ads inline, lead score + AI profile, multilingual contact-page crawl (10+ languages)
  - **New loud PAY\_PER\_EVENT transparency callout** with 🛡️ "Run failed before any lead was delivered? You pay `$0`" — directly addresses the #1 Store-reviewer deal-breaker (hidden charges / timeouts that still cost money)
  - **New "How much will it cost to scrape {city}?" SEO section** with exact formula `Total = (places × $0.003) + (websites × $0.002)` and 9-row cost-per-city table (Budapest / Berlin / London / NYC / Chicago / LA / Sydney). Targets long-tail SEO queries.
  - **Input-fields table expanded** with `autoExtend`, `validateEmails`, `extractWebSignals`, `enrichMetaAds`, `sinceDatasetId` — previously buried in INPUT\_SCHEMA only.
  - `maxResults` default updated from `100` to `10` in the table to match the fast-first-run v1.2.11 change.

#### Infrastructure / discovery signals

- Apify Store search shows the actor as `notice: UNDER_MAINTENANCE` (QA flag from 1.2.11 email). Next QA cycle against v1.2.13 (maxResults.default=10, 113s real runtime vs 5-min QA budget = 62% margin) should clear the flag within ~24 h.
- Apify Store lists `currentPricingInfo: {pricingModel: "FREE"}` — this is because PAY\_PER\_EVENT activation requires a **manual Console submission** (actor.json alone is insufficient). The 14-day notice period kicks in only after Console submission. Flagged for user action; docs + action steps included in the session report.

### \[1.2.12] — 2026-04-23

#### Added

- **Zero-touch timeout UX — `autoExtend` input flag (default `true`)**. The user no longer has to know the "timeout" concept exists. When enabled and the preflight detects the run timeout is too short for `maxResults`:
  1. The actor uses `apify-client` to start a **fresh run** of itself with the same input but a sufficient timeout (`estRuntime × 1.3`, rounded up to the next hour).
  2. The original run exits **SUCCEEDED** with a status message pointing to the new run's dataset URL.
  3. The spawned run receives `autoExtend: false` to prevent cascading loops.
  4. Billing is unchanged — PAY\_PER\_EVENT charges only for leads actually delivered, regardless of which run delivered them.
- When `autoExtend: false` OR the spawn API call fails, falls back to the existing v1.2.9 preflight refusal with the detailed 4-step error message (now prefixed with `[0] ⭐ EASIEST: set autoExtend: true`).

#### Changed

- **`defaultRunOptions.timeoutSecs`: 7200 → 14400 (2h → 4h)** via platform API PUT. The 4-hour default covers every possible `maxResults` up to our 1000 cap with massive headroom (worst-case full-enrichment = ~1 hour real runtime). Combined with `autoExtend`, the user literally never needs to see or touch a timeout field.
- **Three-layer defense in depth** now complete:
  - Layer 1: 4h default covers 99% of users on the "just click Start" path.
  - Layer 2: `autoExtend: true` (default) handles the 1% where timeout was manually lowered (e.g., Apify QA's forced 5-min tests).
  - Layer 3: Preflight refusal as final safety net if `autoExtend: false` or the spawn API fails. Even then, 0 events billed.

### \[1.2.11] — 2026-04-23

#### Fixed

- **Apify automated QA was flagging the actor "Under maintenance"**. Email received 2026-04-23 from Apify: *"your Actor did not pass our automated quality assurance tests during the last three days"*. Apify's QA system runs every Store actor on the prefill input with a **5-minute timeout**. Our prefill had `maxResults: 100` + all enrichments on — estimated 3.2 min, but real-world Google Maps scroll + detail-page loads + slow third-party websites pushed actual runtime over 5 min in 7 of 27 external runs (26% TIMED-OUT). Public stats confirmed: `publicActorRunStats30Days: {SUCCEEDED:18, TIMED-OUT:7, FAILED:1, ABORTED:1}`.
- **`maxResults.default`: 100 → 10.** At 10 leads × 8s × ÷5 concurrency + 30s overhead ≈ 46s — fits comfortably in QA's 5-min window with huge margin for slow websites. Power users wanting 100–1000 just type a new value in the input form; the preflight check (v1.2.7+) protects them from setting it too high for their timeout.
- **Description rewritten** to frame `10` as a "fast first-run for evaluation" and explicitly mention bumping to 100–1000 for production. Clarified timeout UI path: `Start new run modal → ⚙ Options → Timeout`.

#### Why this matters

- "Under maintenance" flag **hides the actor from Store search**, killing organic discovery. Restoring passing QA is mission-critical — the 3 external users/week discovery rate vanishes while flagged. Next QA cycle (within ~24h of push) should lift the flag automatically.

### \[1.2.9] — 2026-04-20

#### Improved

- **Preflight error UX — user-language, explicit "YOUR input"**. First field trial of v1.2.7/8 surfaced that the log said *"Estimated runtime exceeds the run timeout"* (passive, technical). Non-technical users might assume the actor is broken. Rewritten to:
  - Open with **"⛔ YOUR INPUT WON'T FIT IN THE RUN TIMEOUT — STOPPED BEFORE ANY CHARGES"** (active voice, ownership on user's settings).
  - Show **user's own numbers** (`maxResults: 500 → needs ~28 min`, `Your run timeout: 5 min`, `Gap: ~23 min`) so they immediately connect the failure to what they entered.
  - Lead with the money reassurance: **"Zero events billed. Your Apify credit is untouched."**
  - 4 labeled fix-paths `[A]/[B]/[C]/[D]` with dynamic values: `Lower from 500 to ≤ 79`, `Raise timeout to 3600s`, `Split into 7 runs of 79`, `Disable validateEmails (−1s/lead)`.
  - Enrichment-disable tips are **conditional** — only shown for flags currently ON.
  - Footer: **"This is a safety check, not a bug."**
- **Status message** (the one line in the dashboard red banner) rewritten from `"Preflight failed: estimated runtime exceeds run timeout"` to a self-contained human sentence: `"Your input (maxResults=500) needs ~28 min but the run timeout is only 5 min. Stopped before charging you — 0 events billed. See log for 3-4 one-click fixes."`

### \[1.2.8] — 2026-04-20

#### Fixed

- **Preflight refusal now marks run as FAILED (not SUCCEEDED)**. v1.2.7 shipped with `Actor.exit(1)` which, counter-intuitively, still resolves the run as `status: SUCCEEDED, exitCode: 0` on Apify. Users would see a "green checkmark" run with zero results — indistinguishable from a genuinely empty dataset. Swapped to `Actor.fail(statusMessage)`. Now:
  - Run list shows red **FAILED** badge.
  - Dashboard header surfaces the status message directly.
  - Aggregate `publicActorRunStats30Days.FAILED` counter separates preflight refusals from real timeouts in our observability.
  - CLI exits with `Error: Actor failed!` instead of `Success: Actor finished`.

### \[1.2.7] — 2026-04-20

#### Added

- **Preflight timeout refusal**. After pushing v1.2.6 with the 7200s default timeout fix, we observed a fresh external TIMED-OUT event at 2026-04-19 20:52 UTC — the fix alone wasn't sufficient. Root cause: users can (1) override the actor timeout per-run, (2) request a `maxResults` so large even 2h isn't enough, (3) pin an older build. The estimator *warned* about this but didn't stop the run.
  - The actor now reads `APIFY_TIMEOUT_AT` env var (platform-set deadline) and compares it to the runtime estimate.
  - If `estRuntime > availableTime`, the actor **exits with code 1 before scraping a single page** — zero `place-scraped` / `website-scraped` events charged.
  - The error message surfaces 4 concrete fixes with the exact numbers: how many seconds to raise the timeout to, what `maxResults` would fit safely, how to split into multiple runs with `sinceDatasetId`, and which enrichment flags to disable.
  - Preflight success also logs a `✅ Budget OK` line so users see explicit confirmation that the run will fit.
- Result: users discover the misconfiguration in 30 seconds instead of burning 1-2 hours on a run that's mathematically guaranteed to fail.

#### Philosophy

- Fast failure > slow degradation. A user who sees a clear preflight error and fixes their input gets value 5 minutes later. A user whose run silently trucks along for 2 hours before timing out leaves permanently. This is worth the small cost of the "I know what I'm doing, just run it" escape-hatch case (which doesn't exist in our v1.2.x user base per the observed data).

### \[1.2.6] — 2026-04-19

#### Fixed

- **TIMED-OUT runs**. Of the first 14 external runs on the Apify Store, 3 hit the default 1-hour timeout (21% abort rate). Root cause: users requesting 500+ leads with full enrichment couldn't complete in 3600s. Fixes:
  - **`.actor/actor.json`** `defaultRunOptions.timeoutSecs` bumped from `3600` to `7200` (2 hours). Covers ~900 leads at default concurrency.
  - **Runtime estimator** at actor startup. Computes expected runtime based on `maxResults` × `searchQueries.length` × feature flags, logs it on startup, and emits a `⚠️` warning if the estimate exceeds the default timeout — telling the user exactly how much to raise it (and why).
  - **INPUT\_SCHEMA `maxResults` description** updated with timeout guidance so users see it before they launch a 1000-lead job.

#### Added

- **Post-run CTA log**. At the end of every successful run (results > 0), the actor prints a clean footer with:
  - Success summary (leads delivered, high-deliverability count, modern-sites count)
  - Two actionable tips (filter by `deliverability: "high"`, use `sinceDatasetId` for next run)
  - ⭐ Rate-the-actor link — no popups, no emails, shown once after value was delivered
- **README header CTA** — concise bookmark / rate-the-actor line at the top.

#### Why these changes

- First-cohort signal (14 external runs, 3 timeouts) was strong enough to warrant a UX fix rather than just a doc update. New estimator means a 2000-lead user now sees "⚠️ estimated 3.5 hours, current timeout 2 hours, raise to 14400s" in the first 5 seconds of the run — not 2 hours into a hung job.

### \[1.2.4] — 2026-04-18

#### Added

- **Multi-vertical demo gallery**. Three public datasets covering distinct industries and countries to demonstrate real-world hit rates across markets:
  - 🗽 NYC Italian restaurants (`M9Bd8gMh4NglVKIbt`) — 64% email, 56% ownerName
  - ☕ London coffee shops (`ROgK5EsNU6UtTSwFl`) — 70% email, 79% FB page IDs
  - 🦷 Berlin dentists (`gI04MuKrfPF4D4Ui8`) — 85% email, 65% high-deliverability
- `marketing/showcase.md` — cross-vertical dataset gallery with metrics breakdown and "how to reproduce" inputs.
- Updated `marketing/blog-post.md` to reflect v1.2 features (deliverability grading, web signals, delta mode, Meta Ad Library).

#### Observation

- Regulated markets (e.g. German medical professionals under GDPR + Heilmittelwerbegesetz) show 1.5-2× higher email hygiene and deliverability grades than consumer-facing verticals. This cross-vertical data is now visible in the README landing.

### \[1.2.3] — 2026-04-18

#### Added

- **Legal & Compliance** section in README. GDPR data classification table, Legitimate Interest Assessment template, jurisdictional references (US, EU, UK, Hungary). No other Google Maps scraper on the Apify Store ships documentation this detailed.

### \[1.2.2] — 2026-04-18

#### Fixed

- **Google redirect URL pollution**. `normaliseWebsite()` now unwraps `https://www.google.com/url?q=<real>&...` wrapper URLs emitted by Google Maps, so downstream phases (email extraction, web signal analysis, email validation) operate on the real business domain. Previously, ~40% of results were scraping google.com instead of the real website — including false `press@google.com` primary emails.

### \[1.2.0] — 2026-04-18

#### Added

- **Email deliverability grading** (`emailValidation` field). Every `primaryEmail` is graded via MX / SPF / DMARC DNS lookups + best-effort SMTP RCPT TO probe. Output: `{mxRecords, hasSpf, hasDmarc, smtpValid, isCatchAll, deliverability}` with grade `"high" / "medium" / "low" / "unknown"`. Agencies and cold-outreach teams can now skip invalid / catch-all addresses before burning sender reputation.
- **Web quality signals** (`webSignals` + `webQuality` fields). Lightweight "Lighthouse without Lighthouse" — extracted from the HTML already fetched for email scraping, so zero extra cost. Outputs: `httpsOnly`, `mobileResponsive`, `pageSizeKb`, `hasFavicon`, `hasOpenGraph`, `hasStructuredData`. Perfect for web-dev agency outreach targeting.
- Input flags: `validateEmails` (default `true`), `extractWebSignals` (default `true`).

#### Implementation notes

- Node built-in `dns/promises` + `net` — no new dependencies.
- SMTP probe gracefully degrades to DNS-only grading when port 25 is blocked by cloud egress firewalls (as on Apify infra). DNS signals alone cover ~80% of the deliverability picture.
- Per-domain DNS cache — repeated probes on the same domain don't re-query.

### \[1.1.1] — 2026-04-18

#### Added

- **Delta mode** (`sinceDatasetId` input). Pass a previous run's dataset ID and the actor skips every place already present (matched by Google Maps placeId). Workflow win for scheduled runs: users pay only for new leads.
- **Meta Ad Library enrichment** (`enrichMetaAds` input). For every business with a Facebook URL, the actor fetches the FB page, extracts the numeric page ID, and builds a targeted Meta Ad Library lookup URL (`view_all_page_id=<ID>` — ads from that specific page, not a noisy keyword search). Click-through reveals if the business is currently running Facebook/Instagram ads — a strong buying-intent signal.
- New output fields: `facebookPageId`, `metaAdLibraryUrl` (always filled — page-specific when pageId extractable, keyword fallback otherwise).

#### Implementation notes

- `placeId` Set with lowercase normalization for robust matching across runs.
- Delta filtering happens BEFORE detail-page enqueue, so no `place-scraped` events fire for skipped duplicates = users pay nothing for already-known leads.

### \[1.0.24] — 2026-04-18

#### Fixed

- **`reviewKeywords` PUA (Private Use Area) leak**. Material Icons ligatures in Unicode range `U+E000-U+F8FF` were slipping through as `"Sort"`, `"All"`, etc. Added browser-side + Node-side filter: `.replace(/[\uE000-\uF8FF]/g, '')`.
- **`ownerName` extraction rate raised from 0% to 40%**. Updated `NW` regex pattern to accept Mc/Mac/Van prefixes and hyphenated compound names. Previously `"McDonald"` was rejected because of the internal capital D; now "Pam Weekes & Connie McDonald" (Levain Bakery) and "Jatee Kearsley" (Je T'aime Patisserie) extract correctly.

### \[1.0.23] and earlier

- Initial feature set: Google Maps scraping, email extraction from business websites (5-source with ranking), phone/WhatsApp, social media links, tech-stack detection (WordPress/Wix/Shopify/React/Analytics), lead scoring 0-100, hidden gem score, growth signal, budget tier inference, AI-ready outreach profile, suggested cold-outreach opener, Cloudflare email decoding, ROT13 deobfuscation, JSON-LD parsing, contact-page crawl in 10 languages, industry/cuisine classification, booking URL detection (OpenTable/Resy/Tock), website language detection.

***

### Versioning scheme

- **Minor** (1.x.0) — new features
- **Patch** (1.x.y) — bug fixes / docs updates
- Apify auto-increments the patch on every `apify push` within the same `actor.json` version.
