# Changelog of Yelp Scraper - Extract Business Data, Contacts & Reviews (`khadinakbar/yelp-scraper-all-in-one`) Actor

- **URL**: https://apify.com/khadinakbar/yelp-scraper-all-in-one/changelog.md
- **Full Actor documentation**: https://apify.com/khadinakbar/yelp-scraper-all-in-one.md

## Changelog

### 1.9 — 2026-08-12

#### Bounded automated quality-test prefill

- Reduce only the active prefill to one business with review enrichment disabled so the five-minute quality test persists one truthful default-dataset row without unnecessary pagination.
- Preserve customer defaults of 50 businesses, review enrichment enabled, and five reviews per business.

### 1.8 — 2026-08-10

#### Provider fallback integrity repair

- Skip a provider row without a business name or independently useful business fields instead of billing an echoed direct URL as an unknown business.
- Normalize fallback ratings outside Yelp's 1–5 range to `null` and empty strings to `null` before persistence.

### 1.7 — 2026-08-10

#### Publication-readiness surface

- Added a bounded, non-public task SEO map of distinct business-listing workflows; launch tasks remain empty pending exact canary evidence and explicit task-publication approval.
- Corrected README provenance guidance so `data_source` accurately distinguishes native extraction from the provider fallback.

### 1.6 — 2026-08-10

#### Data-integrity and billing repair

- Reject Yelp error documents before they can be emitted as a business solely because an error heading matched the business-name selector.
- Couple each persisted business row to the `yelp-business` PPE event through the SDK; a rejected complete record is never replaced with a smaller billable placeholder.

### 1.5 — 2026-06-11

#### Input schema UI cleanup

- Removed duplicate sectionCaption keys from `location`, `maxReviewsPerBusiness`, and `serpApiKey` so the Apify form no longer renders repeated section headers (Apify only honors sectionCaption on the FIRST property in each section).

### 1.4 — 2026-05-30

#### SerpApi fallback for reliability under Yelp blocking

- When the live scrape returns nothing in search mode (Yelp edge/IP block), the actor now
  automatically falls back to SerpApi so it still returns data instead of failing
- Fallback uses SerpApi's Yelp Search (business enumeration) + Yelp Reviews engines. It
  populates business_name, yelp_url, yelp_business_id, categories, description, and reviews.
  SerpApi's current Yelp parser does NOT expose aggregate rating/review_count/price/phone/
  address, so those are null on fallback records
- Every record now carries a `data_source` field ('scrape' = full live data, 'serpapi' = leaner fallback)
- New inputs: `useSerpApiFallback` (default true) and `serpApiKey` (optional BYOK, secret).
  The actor's own SerpApi key is stored as an encrypted Apify secret (env var), never in code
- Honest-fail still applies: a run fails only if BOTH the scrape and the SerpApi fallback
  return nothing. Direct-URL mode can't use SerpApi (no slug→place_id resolution)

### 1.3 — 2026-05-30

#### Reverted Camoufox — back to Chrome (kept all correctness fixes)

- Testing proved Yelp's 403 is edge/IP-reputation level (blocked even direct /biz/ URLs
  under Camoufox), not a browser-fingerprint problem — so Camoufox added a 713MB image
  and slower Firefox for no anti-bot benefit. Reverted the engine to Chrome.
- All 1.1 correctness/billing/schema fixes remain in place (soft-fail, honest-fail,
  concurrency-safe billing, safePushData, cost-cap, JSON-LD numeric coercion).
- Kept the Chrome path on Crawlee's fingerprint generator + retryOnBlocked + single-block
  session retirement (strictly better than the original 1.0.x hand-rolled masking).
- Reliability now depends on Yelp's IP block easing; durable fix tracked in CLAUDE.md
  (Web Unlocker proxy or Yelp Fusion API).

### 1.2 — 2026-05-30

#### Anti-bot: switched to Camoufox (Firefox stealth browser)

- Replaced Chrome + Crawlee fingerprints with Camoufox — a hardened Firefox build with
  native anti-fingerprinting — to beat Yelp's aggressive 403 blocking
- Per-launch Camoufox config binds spoofed geolocation/timezone/locale to the rotating
  Apify residential proxy exit IP (coherent proxy↔fingerprint↔geo), WebRTC blocked,
  humanized cursor
- Dockerfile now uses the Firefox base image and fetches the Camoufox binary at build
- Firefox is slower: concurrency 3→2, request timeout 60→90s, navigation timeout 60s

### 1.1 — 2026-05-30

#### Reliability & billing hardening (no output-shape change)

- Soft-fail on invalid input (missing query / bad maxResults) — exits cleanly instead of marking the run FAILED, protecting the success-rate metric
- Honest-fail when every request is blocked — no more silent SUCCEED with an empty dataset; returns an actionable "switch to residential proxy" message
- Concurrency-safe billing: reserve-then-fill slot counter prevents over-charging past maxResults under parallel requests
- `safePushData` wrapper — a single schema-invalid record no longer crashes the whole batch
- Upfront cost-cap shown in logs and status before any charge; per-business billed total surfaced live
- Charge only fires after a confirmed dataset push; blocked/failed requests are never billed
- Added `pay_per_event.json` to the repo; wired key-value store schema and `meta.generatedBy` in actor.json
- Run summary (scraped, charged, cost, requests, duration) written to KV `OUTPUT`

### 1.0.0 — 2026-04-13

#### Initial Release

- Extract complete business data from Yelp search results and direct business URLs
- Dual-input mode: search by keyword + location, or scrape specific Yelp business URLs
- Extracts: name, rating, review count, price range, categories, full address (street, city, state, ZIP), phone, decoded website URL, business hours, claimed status, description, photos count, GPS coordinates
- Optional review extraction: up to 20 reviews per business (author, rating, date, text)
- PAY_PER_EVENT pricing at $0.003 per business scraped
- Automatic pagination through Yelp search results
- Session pool rotation for reliability at scale
- Resource blocking for faster scraping
- CAPTCHA detection with auto-retry via fresh session
