# Changelog of Millesima Wine Scraper - Prices & Critic Ratings (`mrbridge/millesima-wine-scraper`) Actor

- **URL**: https://apify.com/mrbridge/millesima-wine-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/mrbridge/millesima-wine-scraper.md

## Changelog

All notable changes to **Millesima Wine Scraper** are documented here.

### v3.1.3 (2026-08-16)

#### Fixed (monetization - paying tiers were stopped short of their own cap)

- The spending guard no longer assumes the FREE price of $0.005 for everyone. The Actor now reads the price the platform will actually bill for this run's user from the SDK charging manager, and asks the SDK how many billable results still fit inside the user's per-run cap before writing each row. The Console grid is tiered ($0.005 FREE down to $0.003 GOLD), so the old flat estimate over-counted every paying tier's spend: a BRONZE user received about 87% of the results they had paid for, a GOLD user about 60%.
- On the platform, a missing or unreadable runtime price now fails the run before the first Millesima request instead of silently falling back to the hardcoded value. The constant survives only as an offline estimate for local development runs.

#### Fixed (pricing manifest drift)

- `.actor/pay_per_event.json` no longer declares a price. The Apify Console monetization settings are the single source of truth for what an event costs, and a local number could only drift away from the tiered grid.

#### Added (graceful shutdown)

- An `aborting` handler is registered before the first request. When the platform asks the run to stop, the Actor stops fetching pages immediately, waits for the single dataset write in flight, flushes the non-billed error rows, then saves a compact `CHECKPOINT` record (current start URL, page, wine index and the number of results actually delivered and billed). The handler is idempotent, never charges manually and never re-pushes a delivered row, so a repeated abort signal cannot double-bill anything. The checkpoint is diagnostic for now and prepares a future resume.
- `OUTPUT` gains `recordsDeliveredAndBilled`, `aborted` and `stopReason`.

#### Changed

- A failure to write a wine to the dataset now fails the run instead of being logged as a scraping error and skipped. An undelivered result is a delivery problem, not a page-level incident.

#### Internal

- The billing counter, the crawl cursor and the abort flags are cleared at the start of every run, so a second run in the same process never inherits the previous one's state. The wine cursor is reset when the crawl moves to a new page, so a checkpoint written during a page fetch cannot point at a wine index belonging to the previous page. Accumulated error rows are also flushed on the fatal path, where the happy-path flush is skipped.
- Billing, delivery and run state moved to `src/charge.ts`, `src/spending-limit.ts` and `src/state.ts`, ported from the sibling wine Actors. Dataset writes are serialized, capacity is checked before each write and the counter only moves after a successful one. Results are still pushed one at a time with no explicit event name: the synthetic `apify-default-dataset-item` event bills every default-dataset row on its own.

***

### v3.1.2 (2026-07-31)

#### Fixed (Store QA - Actor was flagged "Under maintenance")

- Apify's automated quality tests run every Actor with the prefilled input and expect success within 5 minutes. Since v3.1.1 removed the caps, the automated run crawled the entire catalog and timed out, which flagged the Actor. `maxItems` is now **prefilled** at 100 in the Console: the value is visible in the field and can be cleared for an unlimited run. An empty or 0 value still means no limit, and API calls that omit `maxItems` keep extracting everything. `startUrls` gets the same prefill as its default so automated tests are deterministic.

***

### v3.1.1 (2026-07-28)

#### Changed (defaults - no more silent truncation)

- `maxItems` and `maxPagesPerCategory` no longer default to 100 and 10. Left empty, both are unlimited: a run on a category URL (e.g. bourgogne.html) now extracts the whole category (3,000+ wines) instead of stopping at 100 wines / 10 pages. The user's Apify per-run spending limit remains the safety net. Set `maxItems` explicitly for small test runs.

***

### v3.1.0 (2026-07-28)

#### Added (coverage - audit finding F8)

- `formats` array: one entry per purchasable format (bottle, case, magnum...) with `label`, `partNumber`, `listPrice`, `offerPrice`, `availability`, `inStock` and `promotion`. This is the per-format breakdown that `priceMin`/`priceMax` summarize.
- `promotion` (boolean): whether the wine is flagged as on promotion.
- `productType` (string): Millesima's product type label.
- `exclusiveCellar` (boolean) and `isSpirit` (boolean) flags.

#### Changed (reliability - audit finding F6)

- A run that extracts 0 wines while errors were recorded now finishes as FAILED with an explanatory status message, instead of the previous silent SUCCEEDED (the v1.1 incident mode). Nothing is billed on such runs: PPE only bills dataset items.
- A listing page that returns HTTP 200 but parses to 0 wines before any wine was seen for that start URL is recorded in the `run-errors` key-value store record as `ZeroWinesOnListingPage`.
- An unexpected mid-run exception now exits with code 1 (FAILED). Previously the cleanup path exited 0 and the run read SUCCEEDED despite the crash.

***

### v3.0.5 (2026-07-28)

#### Fixed (SEV-4 resilience - retries never stopped)

- HTTP retry backoff now respects got's computed retry rules. The previous custom `calculateDelay` ignored `computedValue` while got defaults `enforceRetryRules` to false, so the `limit: 3` cap and the transient-status list were never enforced: every failed request (HTTP 403 included) retried until the run timeout. This is the likely cause of TIMED-OUT runs on large or proxied inputs. Audit 2026-07-28 finding F1.

#### Fixed (data quality)

- `imageUrl` now returns a full CDN URL (`https://static.millesima.com/s3/attachements/h280px/...`) instead of a bare filename. Audit finding F4.

#### Changed

- Default run timeout raised from 1800s to 3600s so full-catalog runs through residential proxies finish instead of timing out. Audit finding F7.
- README output documentation updated to the v3.0 schema (`priceMin`/`priceMax` + the 11 fields added in v3.0.0). The old `price` field was still documented. Audit finding F5.

#### Internal

- Deploy now goes through an allowlist staging step: only `dist/main.js`, `.actor/`, `package*.json`, `README.md` and `CHANGELOG.md` are uploaded. TypeScript sources, tests and source maps no longer ship. Audit finding F3.
- Prepush script rebuilds the bundle before every push so a stale committed `dist/` can never deploy. Audit finding F2.

***

### v3.0.4 (2026-06-02)

#### Changed (simplify - DRY consolidation, no behavior change)

- Adopted shared `wine-core@0.5.0` `summarizeQuality` + `createErrorSink` (replaces the inline null-rate wiring + the hand-rolled run-errors KV accumulator/flush). Identical behavior, less duplication.

***

### v3.0.3 (2026-06-02)

#### Added (SEV-3 observability - selector-drift early warning)

- New OUTPUT KV record with `fieldNullRates` (per-field null-rate % over a bounded sample of pushed wines) and `driftWarnings`. When a drift-prone field (`name`, `priceMax`, `region`) exceeds 80% null on ≥10 wines, the run logs a loud "possible selector drift" warning - catching Millesima `__NEXT_DATA__`/markup changes before users do. Uses shared `wine-core@0.4.0` `computeNullRates`/`detectDrift`. Audit 2026-06-01 finding X-2.

***

### v3.0.2 (2026-06-02)

#### Fixed (SEV-4 billing - error rows were auto-charged)

- FetchFailed / parse-error / push-error rows are no longer written to the default dataset (where the synthetic `apify-default-dataset-item` event auto-billed them at $0.005 each). They are accumulated and flushed to the KV store key `run-errors` at run end. The spending-limit estimator now counts only `winesPushed` (was `winesPushed + errorCount`, which over-counted and could stop a run early on the user's behalf). Infrastructure/parse failures deliver zero value and must not be billed (Apify PPE doctrine). Audit 2026-06-01 finding X-1.

***

### v3.0.0 (2026-06-01)

#### Changed (BREAKING - schema)

- `price` field split into `priceMin` (cheapest format, typically per-bottle) and `priceMax` (most expensive, typically per-case). The previous unlabelled `price` was the *case* price for multi-format wines (~80% of catalog), conflating bottles and cases for downstream consumers. v3.0 surfaces both honestly.

#### Added

- 11 new fields extracted from the listing page's `__NEXT_DATA__` JSON blob: `color`, `alcohol`, `classification`, `farming`, `subRegion`, `country`, `imageUrl`, `inStock`, `availability`, `description`, `partNumber`.
- 10 additional critic note keys (full 17-critic grid, vs ~7 reliably matched via HTML before).

#### Fixed (SEV-4)

- HTTP retries enabled on transient errors (408, 429, 500-504) with exponential backoff (2s, 4s, 8s). Previously a single 403/5xx killed the entire start-URL pagination.
- Per-page error handling changed from `break` to `continue`: a bad page no longer abandons the entire category.
- Per-card try/catch added in HTML fallback parser to protect a single malformed card from breaking the whole page.

#### Internal

- New `parseListingJson` extractor (primary path, ~80% faster than cheerio cardwalk).
- HTML cheerio parser kept as defensive fallback when `__NEXT_DATA__` is missing.
- Migration note: consumers reading `record.price` must update to `record.priceMin` (per-bottle) or `record.priceMax` (per-case) depending on use case.

### v2.1.0 (2026-05-31)

#### Changed

- Adopted `wine-core@0.3.0`'s `createSpendingLimitChecker(pricePerItem)` helper. The local `isSpendingLimitReached()` inline function in `src/run.ts` was replaced with a one-liner that delegates to the shared helper (consolidation, no behavior change).

### v2.0.0 (2026-05-29)

#### Changed (major)

- **Full rewrite from Python to TypeScript.** Source converted from Python 3.11 + httpx + BeautifulSoup/lxml to TypeScript + fetch (got-scraping) + cheerio. Behavior preserved end-to-end: same selectors, same field extraction, same 24-critic mapping, same pagination logic, same RESIDENTIAL proxy coercion.
- Runtime: `apify/actor-node:22` (was `apify/actor-python:3.11`).
- Package: `wine-core` workspace dependency adopted for `sleep` helper (and ready for future shared logic).
- Build: single-stage Dockerfile with `dist/` produced by `tsup` (ESM bundled).
- Tests: pytest -> vitest, 21 unit tests on `parser.ts` against the same `.html` fixtures (Bourgogne + Bordeaux p2), including the tripwire test for Louis Latour Aloxe-Corton 2018 and the JR -> Jancis Robinson critic-mapping safeguard.
- Atomic charge: `Actor.pushData(item, 'apify-default-dataset-item')` (PPE-006 compliant) replaces the separate push+estimate pattern, with `eventChargeLimitReached` honored to stop early.

#### Migration notes

- The Python source (`requirements.txt`, `src/*.py`, `tests/*.py`, `tests/conftest.py`) was removed from the repo. Git history preserves the legacy implementation.
- Output schema (`dataset_schema.json`, `input_schema.json`) is unchanged: existing consumers see no row-shape difference.
- PPE event (`apify-default-dataset-item` at $0.005) is unchanged.

### v1.2.9 (2026-05-18)

#### Changed

- Store description, SEO description, README intro and comparison table: "8 publications" replaced with "18 publications" (consistent with the v1.2 critic-coverage expansion)
- Cross-promo table: "8 major critics" → "18 major critics"; broken sister-link `wine-searcher-region-scraper` replaced with `wine-searcher-grape-scraper`; "Wine Searcher" → "Wine-Searcher" (hyphenated)
- Related Actors section: added `wine-searcher-scraper-from-list`, replaced broken Wine-Searcher Region link
- Pricing: Starter plan tier renamed `$49 → $29/month` with recomputed capacity (~5,800 wines/month)
- README L13 features bullet: removed `color` field reference (dropped in v1.2), added `producer`
- Issues URL canonical: `apify.com/.../issues` → `console.apify.com/actors/YYFcotMx2QaXrWa0r/issues`
- Tips bullets reformatted: `**Label** - desc` → `**Label**: desc` (de-AI-ification doctrine)
- Removed all em-dashes (zero em-dash policy 2026-05-18)
- Removed 2× `etc.` from data table (regions, critic scores) and 1× from output schema
- Bumped `actor.json version` to `1.2.9`

### v1.2.8 (2026-05-12)

- Defensive proxy-group coercion: if Apify Proxy is enabled with no group selected, the Actor automatically uses the RESIDENTIAL pool (the only one Millesima accepts). Datacenter and AUTO pools are auto-upgraded with a clear warning log. Fixes the "out-of-the-box 403" issue when users left the Console proxy toggle on.

### v1.2 (2026-05-12)

- Rewrote parser for Millesima's Next.js redesign (legacy `.misds-card` selectors no longer match)
- Fixed critic mapping: `JR` is **Jancis Robinson** (was incorrectly labelled Jeb Dunnuck, which is now `JD`)
- Expanded critic coverage to 18 publications including Bettane & Desseauve, Le Figaro, RVF, Alexandre Ma, The Wine Independent, Jean-Marc Quarin, Wine Decider, René Gabriel
- Fixed `VG` mapping: was generic "Vinous", now correctly resolves to **Antonio Galloni** (Vinous)
- Added `producer` field; dropped `color` (no longer available on listing pages)
- Pagination now uses `?page=N`; listing data is self-sufficient so product-detail fetches are no longer needed
- Wired up `proxyConfiguration` (was previously declared but unused); proxy disabled by default because Millesima rejects Apify Proxy datacenter IPs
- Added regression test suite locked against a saved fixture

### v1.1 (2026-03-20)

- SEO-optimized headings, anti-friction messaging

### v1.0.14 (2026-03)

- Improved error handling with graceful exits
- Added per-item try/catch for partial results
- Updated related actors and documentation

### v1.0 (2025-12)

- Initial release with critic ratings extraction from 8 publications
- Pay-per-event pricing at $0.005/wine
