# Changelog of Yelp Scraper (`tri_angle/yelp-scraper`) Actor

- **URL**: https://apify.com/tri_angle/yelp-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/tri_angle/yelp-scraper.md

### 2026-08-03

*Breaking*

- Removed the `since` field. Yelp no longer exposes `yearJoined` anywhere — not on the business
  page, not via `/biz/{id}/props`, and not via GraphQL, on US or Japanese businesses. It had been
  returning an empty string for some time, so no run was producing a value for it.

*Fixes*

- Fixed roughly a third of business pages failing to parse. Yelp serves two different page
  shells and picks one at random per request; only one of them was supported. On the other, the
  scraper threw and burned a retry, and six fields (`amenitiesAndMore`, `operationHours`,
  `upcomingSpecialHours`, `healthScore`, `claimed`, `verified`) would have come back empty.
- Restored `alternateNames`, which had been returning `[]` for every business since Yelp changed
  its page state layout. Mainly affects businesses in non-Latin-script markets, which is what the
  field is for.
- `scrapeStartedAt` is now present in the output. It was being recorded internally but never
  reached the pushed record.
- `maxImages` above ~30 now works. Photo pagination relied on a "next page" link that no longer
  exists in Yelp's markup, and passed an offset Yelp rounds down to a multiple of 30, so every
  request after the first re-fetched the same page.
- Providing `searchTerms` without `locations` now fails with an explanatory error instead of
  finishing successfully having scraped nothing.

*Optimizations*

- Replaced the stealth browser with plain HTTP requests, and residential proxies with datacenter.
  `yelp.com` is blocked at the domain level regardless of proxy tier, while `yelp.com.au` serves
  the same data over datacenter IPs — so the browser and the residential traffic were both being
  paid for without benefit. Runs are substantially faster and cheaper.
- Reduced requests per business from four to three by reading data the page already embeds
  instead of re-fetching it, and cut the GraphQL calls from three operations to one.
- `maxImages` now defaults to 20, which the business page always includes at no extra cost (it
  embeds 22-23). Higher values still work and fetch additional pages as before.
- A GraphQL failure now returns the business with `priceRange: null` instead of dropping the
  record entirely.

### 2024-05-29 (v0.0.66)

*Fixes*

- Decode html entities (such as `&amp;` → `&`) in place names

### 2024-04-25 (v0.0.64)

*Other*

- Removed `debugLog` and `maxRequestRetries` input options
- Removed `scrapeReviewerName` and `scrapeReviewerUrl` input options, and add this data to the output by default

#### 2024-02-26 (v0.0.62)

*Fixes*

- Adjust to new method of extracting reviews (the Yelp website changed and the old method stopped working)

#### 2023-11-22 (v0.0.61)

*Features*

- Added field `ownerReplies` to each review. This field is an array of objects `{"text": "Review response text", "date": "1/1/2023"}` for each response to the review.

#### 2023-11-08 (v0.0.60)

*Fixes*

- Only provide field `cuisine` for restaurants
- Improved extraction of categories

#### 2023-09-01 (v0.0.59)

*Features*

- New field: `aboutTheBusiness`: extracts info about the business provided by the owner, specifically text about *Specialties* and *History* and year of establishment
- New input option: `debugLog` (disabled by default)

*Optimizations*

- If the scraper is started with `reviewLimit: 0`, the scraper will now completely skip making requests for the reviews page, saving a small bit of time and data transfer

#### 2023-03-20

*Features*

- Add field `alternateNames` to output: provides alternative names, especially useful for places in countries using non-latin characters.

#### 2023-03-01

*Features*

- Rewrite of the scraper to use the new SDK V3

#### 2023-01-12

*Features*

- Allow users to specify `reviewsLanguage` input field, which is the language in which the reviews should be scraped (Only the reviews in the selected language will be scraped).
- Added `availableReviewsLanguages` field to output, which contains a list of languages in which the business has reviews in.

#### 2021-11-19

*Features*

- Update SDK

*Fixes*

- Random crash when scraping images
- Fix log for number of scraped businesses
- Fix website extraction

#### 2021-04-19

*Fixes*

- Fixed page layout change (whole scraper was broken)

#### 2021-03-30

- Added support to different languages domains such as yelp.fr to input url.

#### 2020-03-25

*Features*

- Enhanced reviews with more fields: `language`, `isFunnyCount`, `isUsefulCount`, `isCoolCount`,`reviewerName`, `reviewerUrl`, `reviewerReviewCount`, `reviewerLocation`
- Scraping `reviewerName`, `reviewerUrl` requires enabling personal data input fields: `scrapeReviewerName`, `scrapeReviewerUrl`
- Added section about GDPR and personal data protection to README

#### 2020-02-20

- `searchTerm` and `location` deprecated in favor of `searchTerms` and `locations`. You can scrape any number of those in a single run now.
- Refactored to SDK 1

#### 2020-12-01

- Fixed for new layout

#### 2020-09-25

- Added `maxRequestRetries` to input and increased its default from `2` to `10`
- Added `cuisine` to output
- Added `website` to output
- Added `images` to output

#### 1.0.0

- Data format changed (refer to README.md)
- Fixed scraping information of business
- Updated SDK to 0.21+
- Minor changes to code style and linting
- Updated dependencies
- `priceRange` field changed from `$` / `$$$` to actual prices like `$10-30`
- Reviews dates are now a proper ISO date time string
- Review texts now contains plain text instead of HTML
