# Changelog of Yandex Realty Scraper (`normdata/yandex-realty-scraper`) Actor

- **URL**: https://apify.com/normdata/yandex-realty-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/normdata/yandex-realty-scraper.md

## Changelog

### 2026-09-12 (9)

- `includePriceHistory` is now its own billed event (`price-history-fetched`, charged once per row
  it's actually fetched for - nothing charged on a failed fetch). Verified directly on two real
  runs of the identical 300-row query: with it on, the run took 5.6x longer (154s vs 27s) and
  downloaded 13x more data (208MB vs 16MB) - a real cost, not something to fold into the base row
  price for free.

### 2026-09-12 (8)

- Progress logging: the per-round "collected so far" line is no longer gated behind a minimum
  category size, now also shows request count and how many price bands are still queued, and the
  per-split detail is at debug level (invisible by default) - a real run showed only two INFO
  lines total before it looked stuck, with nothing to show it was still working.
  - Tested `BAND_CONCURRENCY` at 10 instead of 5 directly: *slower* (67s vs 59s on the same
    200-row query), not faster - each page here is a 3-5MB HTML document, so more concurrent
    fetches compete for bandwidth rather than adding real throughput. Left at 5.

### 2026-09-12 (7)

- Fixed a real bottleneck reported from an actual Apify run: the price-range splitting loop
  probed one band at a time, fully serially - on a broad/unfiltered search (tens of thousands of
  listings, dozens of splits needed) that meant paying every single request's latency back to
  back with no overlap. Now processes several bands (5 at a time) concurrently per round, the same
  way page-by-page pagination already did. Verified: 300 rows against the same unfiltered
  61,000+-listing category that was slow before, in 97s, still 300/300 unique.

### 2026-09-12 (6)

- Added `includePriceHistory` (search/url modes): fetches each result's price history in the same
  run instead of requiring a separate `lookup` call per offer ID afterwards - real UX friction a
  user pointed out. Off by default since it costs one real extra request per row.

### 2026-09-12 (5)

- Fixed the `roomsTotal` input field, which required typing raw JSON (e.g. `[2, 3]`) - now a
  proper multi-select checkbox list (1, 2, 3, 4, 5+), matching the rest of the portfolio's
  convention for small fixed-option fields.

### 2026-09-12 (4)

- Found (via a real 25-page sweep of a single filtered search) that realty.yandex.ru's own search
  index can briefly lag a listing's live price: about 2-3% of rows in a filtered search can have a
  live price outside the requested `priceMin`/`priceMax` at the moment this Actor reads them,
  because the listing's own price changed between when Yandex indexed it as a match and when the
  full listing was fetched. Confirmed directly on one real listing whose price swung between 22.4M
  and 26.4M RUB five times in six days - not a clustering artifact (every case checked had
  `clusterSize: 1`), not this Actor's own bug (a fresh single request never showed a violation;
  only a slower sweep across many requests did). Added `price_out_of_requested_range`: rather than
  silently dropping the row (which would waste the fetch this run already paid for in compute with
  nothing gained - no request is saved either way) or silently keeping it unexplained, it's kept,
  billed normally, and flagged so it's clearly not a bug.

### 2026-09-12 (3)

- Fixed a real bug found on the very first live Apify run: a small `maxItems` (e.g. the default
  10-row prefill) against a huge, unfiltered category still triggered the price-range splitting
  machinery before ever checking whether the first page alone already had enough rows - a run
  asking for 10 rows against 61,000+ matching listings took over a minute and had to be aborted.
  Fixed by always collecting whatever a page already returned *before* deciding whether to split
  or paginate further; a small `maxItems` request now short-circuits after a single page fetch.
  Verified: 10 rows against the same unfiltered 61,000+-listing category dropped from 60+s (hung)
  to ~4s.
- Fixed the Apify input schema failing to build (`dealType`, `category`, `areaMin`, `areaMax` were
  missing the `description` field Apify requires on every input property).

### 2026-09-12 (2)

- Verified the full `search` + splitting path at 1,000 rows against a genuinely huge, unfiltered
  category (61,500+ Moscow apartments for sale): **1,000/1,000 unique `offer_id`, 0 duplicates**,
  after 60 real price-range splits plus 54 bands that hit the split/request budget and fell back
  to their first page - field fill rates matched the earlier 500-row run exactly, no degradation
  from splitting. Found and fixed two real bugs surfaced only at this scale (not visible at
  100-500 rows):
  - The per-range probe fetch inside the split loop wasn't wrapped in try/catch - a single timed-
    out request crashed the entire run instead of being skipped like any other unreachable band.
  - A crash of any kind lost every row already collected, because `Actor.fail()` was called
    without first pushing what had been gathered - the same class of bug the `aborting` handler
    already covers for a cancelled run, just for an uncaught exception instead. Both are now
    handled: a failed probe is skipped and counted as an incomplete band, and the top-level catch
    pushes whatever's in memory before failing.
  - Splitting on the arithmetic midpoint of a price range converges far slower than expected on a
    real, skewed price distribution (most listings cluster well below a market's own headline
    max). Switched to splitting on the *median of the real prices already sampled* from that
    range's own first page (free - those rows are fetched either way), which roughly halves the
    matching listings instead of the numeric range.
  - Bumped the per-page timeout from 25s to 45s after measuring live page-fetch times ranging from
    \~2s to over 40s under different real conditions.

### 2026-09-12

- Initial release: `search` (city/deal-type/category/rooms/price/area filters), `url` (paste any
  realty.yandex.ru search URL), and `lookup` (specific offer IDs) modes over realty.yandex.ru's own
  server-rendered pages - no API key, no login, no browser, no captcha hit at any point.
- One flat row per listing: price (+ trend and previous price), area/rooms/floor, exact address
  with GPS and metro walk time, building/renovation detail, seller info, full description, photos.
- `lookup` mode additionally returns a real dated price-history series per offer.
- Found and fixed via testing at 100/500/1,000+ rows before shipping:
  - realty.yandex.ru caps every filtered search at 25 reachable pages (500 rows) regardless of how
    many listings actually match - confirmed live, not assumed. Worked around with automatic price-
    range splitting (bisect and search each half independently) so a broad search can still reach a
    large `maxItems`, bounded by a real request/split budget so a run can't spiral on a market where
    many listings share round prices.
  - `currency` is really `"RUR"`, not `"RUB"` as first assumed - caught by a unit test against a
    real fixture value before shipping, not by manual inspection.
  - `url` isn't consistently absolute in the source (protocol-relative for a Yandex-native listing,
    already-absolute for one syndicated from a partner site) - normalized to `https://` either way.
- Phone/contact extraction is not implemented - revealing a number is a client-side action behind
  what looks like a session/token flow this Actor doesn't attempt; documented honestly rather than
  guessed at.
- Free-plan runs are locked to a fixed 10-row sample.
