Get price, currency, stock, GTIN (UPC/EAN), SKU, MPN, brand, name and image from any shop's product pages, read from their schema.org Product data. Paste product URLs or whole stores (Shopify, WooCommerce and more, via sitemaps). Monitor mode returns only changed prices and stock.
Versions follow MAJOR.MINOR.PATCH (src/version.py); Apify shows MAJOR.MINOR from .actor/actor.json.
Every run logs its version and records it in the RUN_STATS key-value record.
1.1.0 (2026-09-29)
Chain it after another actor, e.g. an e-commerce scraper or Sitemap URL Extractor: price and stock of every product
page it found.
New inputs Or: product URLs from a dataset (datasetId, a dataset picker with read access) and Field with
the product URL (datasetUrlField). With a dataset, each item's link is read like a line of productUrls, and
productUrls and stores are ignored: productUrls has a default that Apify fills into API and integration runs,
so a chained run never fetches or charges for the Wikipedia Store examples, and a store left in a copied input
isn't read either (the log says so). URL patterns and Max pages per store stay store-only. In an integration,
"datasetId": "{{resource.defaultDatasetId}}" passes the finished run's results. The field is found automatically
when left empty: the first of productUrl, url, link, loadedUrl with a web address in the first 100 items,
then the same names one level down (product.url); a dotted path names a nested one. A Google Maps link is never
used.
Only changed prices and stock keeps its memory key (the saved task and the URL patterns): the dataset isn't
part of it, so a chained run, which reads a new dataset every time, compares with the last run, and a product URL
from a dataset shares its memory with the same URL typed in (the product id is the same). Inputs without a
dataset keep their memory from earlier runs.
Items without a link and values that aren't a public web address are skipped, never fetched or charged; a product
listed twice is read and charged once; tracking parameters (utm_*, gclid, fbclid, ...) are dropped. All
counted in the new RUN_STATS.dataset record (itemsRead, websites (the product URLs kept), withoutWebsite,
notAWebsite, duplicates, the field used and whether it was found automatically). A dataset with no product URL
at all, or one that can't be read, fails the run with the reason.
New nullable fields on every row: sourceTitle (the item's title, else name) and sourceIndex (its position
in the dataset, 0 = the first), to join products back to the items; null for products not read from a dataset.
Read with the run's own token (the user's access, read-only), only the fields needed; at most the first 20,000
items and 10,000 product URLs per run.
No change to the price, the charged events (unchanged-product included), or runs without a dataset.
1.0.0 (2026-09-28)
First release.
Reads a product page's own schema.org data: JSON-LD first (the shared mms_common.jsonld reader, lenient about
raw newlines, CDATA/comment wrappers and trailing semicolons), then microdata. One row per product page: name,
brand, SKU, MPN, GTIN, image, price, low/high price, currency, availability, condition, offer count, up to 50
offers (variants), rating and review count. Nothing is taken from the visible text.
Offer normalisation: a single Offer, a list of offers, AggregateOffer (with or without offers inside),
priceSpecification (strikethrough/list prices skipped), ProductGroup + hasVariant, @id references.
Prices as numbers from the usual spellings ("1,299.00", "1.299,00", "19,99", "$19.99"); ranges and negatives are
null. Currency only as an ISO 4217 code. Availability and condition as plain schema.org names from URLs, prefixes
or loose spellings. GTIN-8/12/13/14 with the check digit verified; failing codes are left out and counted
(RUN_STATS.rejectedGtins). The page's main product: the one whose URL is the page's, else the first product the
page states as its own; a page listing several products (an ItemList) and none of its own is a listing page.
Whole-store mode: product pages from the shop's product sitemaps (Shopify sitemap_products_*.xml,
WooCommerce/Yoast product-sitemap.xml, WordPress wp-sitemap-posts-product-*.xml, BigCommerce
xmlsitemap.php?type=products), found through robots.txt Sitemap: lines or /sitemap.xml, via the shared
mms_sitemap. Localised copies next to an unprefixed one are skipped. URL patterns, and Max pages per store
(default 100, up to 10,000). A store without product sitemaps is read in its sitemap's order.
Only changed prices and stock since the last run (onlyChangedProducts): per search (URL patterns) and saved
task, one gzipped memory record in the user's own account. Rows only for new products and changed price, price
range, currency, availability or offers, with changes, previousPrice, previousAvailability,
previousCheckedAt. Unchanged products are charged the unchanged-product event. Pages that fail, are blocked or
state no product keep their memory and are never reported as a change; only delivered products update it.
Never charged, and listed in the NOT_RETURNED record with a reason: bot checks served with 200 and 401/403/451
answers (blocked; never worked around), pages robots.txt disallows (robots), pages without product data
(no-product-data), listing pages (listing-page), 404/410, non-HTML files and failures. A shop that blocks 3
pages in a row isn't asked again in that run; a store whose first 25 pages state no product is stopped.
Polite: the shared mms_common client (honest User-Agent, robots.txt and Crawl-delay, private-network guard, ports
80/443), at most 2 pages in flight and 1 page start a second per shop (a longer Crawl-delay wins), 10 pages in
flight in all. No proxies of any kind.
Charged per product returned (Apify's apify-default-dataset-item event); unchanged-product in the monitoring
mode. Max results per run and the maximum cost per run are honoured before a page is read. A run fails only when
every input was broken (typo, dead page); shops that block or publish no product data don't fail the run.
Source and terms (read 2026-09-28): the actor reads pages the user names, and only the structured data those pages
publish for machines (schema.org, the markup shops add for search engines). The default input reads three product
pages of the Wikipedia Store (store.wikimedia.org, run by Shopify for the Wikimedia Foundation). Its robots.txt says
"Public product, collection, page, blog, policy, cart, and localized HTML is crawlable" and allows /products/. Its
terms of service (https://store.wikimedia.org/policies/terms-of-service ) have no clause on automated access; they
reserve the store's text and images ("Nothing on this site grants a license or right to use Wikimedia Foundation
trademark or copyright protected material"), so the output carries facts (price, stock, identifiers), the product's
name and links, never the shop's descriptions or images themselves.