# Changelog of Shopify Catalog · Products by Store URL (`corent1robert/shopify-products-scraper`) Actor

- **URL**: https://apify.com/corent1robert/shopify-products-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/corent1robert/shopify-products-scraper.md

## Changelog

All notable changes to this Actor are documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/).

### \[1.8] - 2026-09-07

#### Added

- **Actor Standby:** HTTP lookup `GET`/`POST /lookup` — **one** Shopify storefront URL, up to **5 products**. Console **Start** still exports the full catalog as before.
- OpenAPI spec for the Standby tab (`web_server_openapi.json`).

#### Changed

- `startUrls` is no longer schema-required so Standby can boot with empty INPUT. Console **Start** still exits if the list is empty.

### \[1.7.3] - 2026-04-09

#### Fixed

- **Console input:** removed duplicate **What to scrape** section header — `sectionCaption` only on `startUrls`; **Product cap** stays in the same block without repeating the title.

#### Added

- **Shopify-only guard:** before scraping each store, the Actor verifies the hostname looks like a **Shopify** storefront (public `products.json` shape, HTML markers such as `cdn.shopify.com`, or `cart.js` token). Otherwise it **throws** with a clear English message — no silent empty dataset from wrong platforms.
- **Console copy + run banner** explicitly state **Shopify URLs only**; README FAQ updated.

### \[1.7.2] - 2026-04-09

#### Changed

- **Console input (UX):** one **What to scrape** block (URLs + product cap together), shorter copy, friendlier labels; second block **Connection (optional)** for proxy only. Top description mentions **per-row pricing** at a glance.

### \[1.7.1] - 2026-04-09

#### Added

- **`resolveProductCap`** moved to `src/lib/inputResolve.js` with **unit tests** (`tests/inputResolve.test.js`) so the “full catalog vs cap” rules are regression-tested.

### \[1.7] - 2026-04-09

#### Changed

- **`maxProducts` default:** **Omit**, empty, **0**, or invalid → **full** public catalog per store. A positive number limits how many products per store. Console copy and schema **no longer prefill 25**; Step 2 is framed as an **optional** cap.

### \[1.6] - 2026-04-09

#### Changed

- **Dataset / output schema:** removed the **Full export** custom view; the Console **All fields** tab already shows every column. **Catalog overview** remains for a slim reporting slice.

### \[1.5] - 2026-04-09

#### Changed

- **Apify Console input** is now **only** `startUrls`, `maxProducts`, and `proxyConfiguration`. Removed: verbose logs, include body HTML, one row per product, enrich from product page — behavior is **fixed in code** (`EXPORT_INCLUDE_BODY_HTML`, `EXPORT_ONE_ROW_PER_PRODUCT`, `EXPORT_ENRICH_FROM_PRODUCT_PAGE`, `RUN_VERBOSE_LOGS` at top of `src/main.js`, all `false` by default). Fork or contact for custom shapes.

### \[1.4] - 2026-04-09

#### Changed

- **Apify Console input**: removed **batch size**, **parallel requests**, **pause between batches**, and **max product sitemap files** from the form. Sensible defaults remain in code; power users can still set `batchSize`, `handleRequestConcurrency`, `handleBatchCooldownMs`, and `maxSitemapIndexes` via the **API** or local `input.json` (see README).
- **Console copy**: slightly shorter descriptions on output toggles; **enrich from product page** no longer names implementation details in the form text.

### \[1.3] - 2026-04-09

#### Added

- **RUN\_LOG / console progress**: ASCII **progress bar**, **percent**, **ETA** (when total is known), and elapsed time for `products.json` pagination (with `maxProducts` cap) and for **sitemap / headless handles**. Uncapped catalog runs show throughput (`prod/s`) and `ETA n/a`. On Apify Cloud, **run status** updates at most every ~12s with the same percentages.

- **`default-input.json`** + **`npm run reset-apify-input`**: after **Purge storages** in Apify CLI / Console local, `storage/.../INPUT.json` can lose `startUrls` (required by the input schema). Run the script to copy the default file back. `storage/key_value_stores/default/INPUT.json` is tracked in git with a valid template.

#### Changed

- Local **output.csv** only includes columns where **at least one row** has a non-empty value; tag slots are limited to the **maximum tag count** in the run (no trailing empty `tag_*`). Dataset fields are unchanged.
- Local **output.csv** is **refreshed during the run** (after each products.json page and each sitemap/handle batch): new rows are **appended** when column headers are unchanged; the file is **rewritten** when new columns appear (e.g. more tag slots). Helps long headless runs and crash recovery.
- **Sitemap fallback**: when `/products/{handle}.json` is blocked (e.g. HTTP 403) or not Shopify-shaped JSON, the Actor tries the **public product page HTML** and parses **Next.js `__NEXT_DATA__`** when it contains **Gymshark-style `productData`** (headless storefronts).
- **CDN throttling (HTTP 405)**: local requests use **browser-like headers**; **429 / 405 / 5xx** are retried with backoff (body drained between tries). Sitemap fallback uses **limited parallelism** (`handleRequestConcurrency`, default 3) and **pause between batches** (`handleBatchCooldownMs`, default 900 ms). Tuning: raise cooldown to **3000–8000** ms and set concurrency to **1–2** on strict storefronts.
- **Network errors**: local runs use **`node:undici` `fetch`** (same as the platform path). Failures now log **`error.cause`** (e.g. `ENOTFOUND`, `ECONNRESET`, certificate issues) instead of only the generic `fetch failed`.
- **HTTP reliability**: every `fetch` response body is **read to completion** (including failed status codes). Skipping `res.text()` on 403 with **parallel** handle requests could reuse connections incorrectly under undici and return **wrong/truncated HTML**, breaking headless PDP parsing on stores like Gymshark.

#### Added

- **Full storefront JSON mapping**: barcode, weight, grams, currencies, inventory/fulfillment flags, quantity rules, variant timestamps, gallery `imageSrcs`, featured image, product options summary, `published_scope`, `template_suffix`, etc.
- Optional **enrich from product page** (`enrichFromProductPage`): HTTP GET on the public PDP, parse **Schema.org JSON-LD** into `pageLd*` columns (description, brand, offer price/currency, availability URI, GTIN, aggregate rating when present).

### \[1.2] - 2026-04-09

#### Changed

- Local **output.csv** no longer embeds JSON: **tags** are split into `tag_01` … `tag_15`; nested **variants** (one-row-per-product mode) add `variant_01_*` … `variant_20_*` columns (omitted when every row is already one variant). Dataset JSON is unchanged.

### \[1.1] - 2026-04-09

#### Changed

- Default **Maximum products per store** is **25** for safe Console and smoke runs; use **0** for a full catalog (see README).
- **startUrls** prefill stays a **single** storefront URL for gentle first runs.
- URLs are **sanitized** (trim + strip line breaks) before parsing.

#### Added

- **output.csv** written on **local** runs (`apify run` / `npm start`), UTF-8 with BOM, `;` separator (Excel-friendly).

### \[1.0] - 2026-04-09

#### Added

- Initial release: export public Shopify storefront catalogs from one or more URLs.
- Plan A: paginated `products.json` collection.
- Plan B: `collections/all/products.json`, then sitemap-driven `products/{handle}.json` fallback.
- Dataset view **Catalog overview** (slim columns); live **RUN\_LOG** in Key-Value Store. Use the Console **All fields** tab for the complete row shape.
- Input options: `maxProducts`, `includeBodyHtml`, `oneRowPerProduct`, `batchSize`, `maxSitemapIndexes`, `proxyConfiguration`, `verboseLogs`.
- Unit tests for pure URL and mapping helpers (`npm test`).
