# Changelog of Product Price & Stock Monitor: any shop's schema.org data (`humble-echidna/product-offers`) Actor

- **URL**: https://apify.com/humble-echidna/product-offers/changelog.md
- **Full Actor documentation**: https://apify.com/humble-echidna/product-offers.md

## Changelog

Versions follow MAJOR.MINOR.PATCH (`src/version.py`); Apify shows MAJOR.MINOR from `.actor/actor.json`.
Every run logs its version and records it in the `RUN_STATS` key-value record.

### 1.1.0 (2026-09-29)

Chain it after another actor, e.g. an e-commerce scraper or Sitemap URL Extractor: price and stock of every product
page it found.

- New inputs **Or: product URLs from a dataset** (`datasetId`, a dataset picker with read access) and **Field with
  the product URL** (`datasetUrlField`). With a dataset, each item's link is read like a line of `productUrls`, and
  `productUrls` and `stores` are ignored: `productUrls` has a default that Apify fills into API and integration runs,
  so a chained run never fetches or charges for the Wikipedia Store examples, and a store left in a copied input
  isn't read either (the log says so). URL patterns and Max pages per store stay store-only. In an integration,
  `"datasetId": "{{resource.defaultDatasetId}}"` passes the finished run's results. The field is found automatically
  when left empty: the first of `productUrl`, `url`, `link`, `loadedUrl` with a web address in the first 100 items,
  then the same names one level down (`product.url`); a dotted path names a nested one. A Google Maps link is never
  used.
- **Only changed prices and stock** keeps its memory key (the saved task and the URL patterns): the dataset isn't
  part of it, so a chained run, which reads a new dataset every time, compares with the last run, and a product URL
  from a dataset shares its memory with the same URL typed in (the product `id` is the same). Inputs without a
  dataset keep their memory from earlier runs.
- Items without a link and values that aren't a public web address are skipped, never fetched or charged; a product
  listed twice is read and charged once; tracking parameters (`utm_*`, `gclid`, `fbclid`, ...) are dropped. All
  counted in the new `RUN_STATS.dataset` record (`itemsRead`, `websites` (the product URLs kept), `withoutWebsite`,
  `notAWebsite`, `duplicates`, the field used and whether it was found automatically). A dataset with no product URL
  at all, or one that can't be read, fails the run with the reason.
- New nullable fields on every row: `sourceTitle` (the item's `title`, else `name`) and `sourceIndex` (its position
  in the dataset, 0 = the first), to join products back to the items; `null` for products not read from a dataset.
- Read with the run's own token (the user's access, read-only), only the fields needed; at most the first 20,000
  items and 10,000 product URLs per run.
- No change to the price, the charged events (`unchanged-product` included), or runs without a dataset.

### 1.0.0 (2026-09-28)

First release.

- Reads a product page's own schema.org data: JSON-LD first (the shared `mms_common.jsonld` reader, lenient about
  raw newlines, CDATA/comment wrappers and trailing semicolons), then microdata. One row per product page: name,
  brand, SKU, MPN, GTIN, image, price, low/high price, currency, availability, condition, offer count, up to 50
  offers (variants), rating and review count. Nothing is taken from the visible text.
- Offer normalisation: a single `Offer`, a list of offers, `AggregateOffer` (with or without offers inside),
  `priceSpecification` (strikethrough/list prices skipped), `ProductGroup` + `hasVariant`, `@id` references.
  Prices as numbers from the usual spellings ("1,299.00", "1.299,00", "19,99", "$19.99"); ranges and negatives are
  null. Currency only as an ISO 4217 code. Availability and condition as plain schema.org names from URLs, prefixes
  or loose spellings. GTIN-8/12/13/14 with the check digit verified; failing codes are left out and counted
  (`RUN_STATS.rejectedGtins`). The page's main product: the one whose URL is the page's, else the first product the
  page states as its own; a page listing several products (an `ItemList`) and none of its own is a listing page.
- Whole-store mode: product pages from the shop's product sitemaps (Shopify `sitemap_products_*.xml`,
  WooCommerce/Yoast `product-sitemap.xml`, WordPress `wp-sitemap-posts-product-*.xml`, BigCommerce
  `xmlsitemap.php?type=products`), found through robots.txt `Sitemap:` lines or /sitemap.xml, via the shared
  `mms_sitemap`. Localised copies next to an unprefixed one are skipped. URL patterns, and Max pages per store
  (default 100, up to 10,000). A store without product sitemaps is read in its sitemap's order.
- **Only changed prices and stock since the last run** (`onlyChangedProducts`): per search (URL patterns) and saved
  task, one gzipped memory record in the user's own account. Rows only for new products and changed price, price
  range, currency, availability or offers, with `changes`, `previousPrice`, `previousAvailability`,
  `previousCheckedAt`. Unchanged products are charged the `unchanged-product` event. Pages that fail, are blocked or
  state no product keep their memory and are never reported as a change; only delivered products update it.
- Never charged, and listed in the `NOT_RETURNED` record with a reason: bot checks served with 200 and 401/403/451
  answers (`blocked`; never worked around), pages robots.txt disallows (`robots`), pages without product data
  (`no-product-data`), listing pages (`listing-page`), 404/410, non-HTML files and failures. A shop that blocks 3
  pages in a row isn't asked again in that run; a store whose first 25 pages state no product is stopped.
- Polite: the shared `mms_common` client (honest User-Agent, robots.txt and Crawl-delay, private-network guard, ports
  80/443), at most 2 pages in flight and 1 page start a second per shop (a longer Crawl-delay wins), 10 pages in
  flight in all. No proxies of any kind.
- Charged per product returned (Apify's `apify-default-dataset-item` event); `unchanged-product` in the monitoring
  mode. Max results per run and the maximum cost per run are honoured before a page is read. A run fails only when
  every input was broken (typo, dead page); shops that block or publish no product data don't fail the run.

Source and terms (read 2026-09-28): the actor reads pages the user names, and only the structured data those pages
publish for machines (schema.org, the markup shops add for search engines). The default input reads three product
pages of the Wikipedia Store (store.wikimedia.org, run by Shopify for the Wikimedia Foundation). Its robots.txt says
"Public product, collection, page, blog, policy, cart, and localized HTML is crawlable" and allows `/products/`. Its
terms of service (https://store.wikimedia.org/policies/terms-of-service) have no clause on automated access; they
reserve the store's text and images ("Nothing on this site grants a license or right to use Wikimedia Foundation
trademark or copyright protected material"), so the output carries facts (price, stock, identifiers), the product's
name and links, never the shop's descriptions or images themselves.
