Product Price & Stock Scraper: Shopify, WooCommerce avatar

Product Price & Stock Scraper: Shopify, WooCommerce

Pricing

from $0.90 / 1,000 product results

Go to Apify Store
Product Price & Stock Scraper: Shopify, WooCommerce

Product Price & Stock Scraper: Shopify, WooCommerce

Scrape product prices and stock from Shopify, WooCommerce, Magento, Shopware, PrestaShop, BigCommerce, Wix and Squarespace shops, not marketplaces. Give product URLs, whole shops or search terms; get one row per product: price, currency, stock, SKU, GTIN/EAN. For e-commerce price monitoring.

Pricing

from $0.90 / 1,000 product results

Rating

0.0

(0)

Developer

Victorix

Victorix

Maintained by Community

Actor stats

1

Bookmarked

2

Total users

1

Monthly active users

6 hours ago

Last modified

Share

What is Product Price & Stock Scraper?

Product Price & Stock Scraper: Shopify, WooCommerce is an e-commerce scraper that reads product prices, stock, SKUs, barcodes (the GTIN, as EAN or UPC) and brands from independent online shops on Shopify, WooCommerce, Magento, Shopware, PrestaShop, BigCommerce, Wix and Squarespace: give it product page URLs, whole shops or search terms, and it returns one row per product; marketplaces such as Amazon and eBay are out of scope. Read live on 2026-09-24, 154 product pages of 46 independent shops on those eight platforms gave a complete row (name, price, ISO currency and stock) for 87.6 % of the pages still online (134 of 153).

Use cases

  • Competitor price monitoring on a schedule: pricing teams and shop owners save the input as a task and run it daily or hourly, without a SaaS monitor's monthly minimum.
  • Back-in-stock and out-of-stock checks: every row says whether the product is in stock; compare two runs to see what came back or ran out.
  • MAP and reseller price checks for brands: read your products on resellers' shops and compare each price with your minimum advertised price.
  • Supplier catalogue export for dropshipping and catalogue teams: a supplier's whole catalogue (prices, stock, SKUs, GTINs) in one shape.
  • Live product facts for AI agents, spreadsheets and databases: the same fields whatever platform a shop runs on, with no scraper to write per site; every row carries a working product URL and the minute it was read.

What you get

One row per product, or per variant on request: name, price (a whole number of minor units and a decimal string), the ISO 4217 currency, the regular price a sale strikes through, in stock or not, SKU, GTIN, MPN, brand, category, attributes, shipping options, image link, rating, the shop's name, the description as plain text and the product URL, exported as JSON, CSV, Excel, XML or HTML.

Price: pay per result, $1.35 per 1,000 products at Bronze ($5.40 on the Free tier) plus $0.000081 a run, and $0.90 per 1,000 rows read through Apify's residential IPs after a plain refusal; blocked, disallowed and data-less pages cost nothing, and so does listing a shop.

Main limit: marketplaces, shops behind a bot wall and storefronts that write their product data with JavaScript are not read; their pages end as free rows.

What data can you scrape from online shops?

The Actor reads what shops publish for search engines and platforms: schema.org Product markup (JSON-LD, microdata, RDFa), Open Graph product tags, Shopify's catalogue JSON and search page, and WooCommerce's Store API. It needs no selectors and nothing breaks when a shop changes its layout.

  • Product pages, or a whole shop. Paste product page URLs from any shop whose pages carry that markup, give a shop's home page, or one of its collection or category pages: Shopify shops are read through their catalogue feed, WooCommerce shops through their Store API, and other shops (Magento, Shopware, PrestaShop, BigCommerce, Wix, Squarespace and more) through the product pages their sitemaps list.
  • Search shops for the products you name. Add search terms beside your shops and each shop is searched for each term instead of read whole: a product matches when every word of the term is a word of its name, or when the term is its SKU or barcode. Shopify and WooCommerce shops are searched through their own catalogue or search, other shops through the addresses of their product pages (best effort).
  • Prices you can compute with: price is an integer in the currency's minor units (3999 is 39.99 EUR, 1000 is 1000 JPY) beside priceDecimal ("39.99"), with the ISO code the shop states. Never a float, never a currency guessed from a symbol.
  • Stock as the shop states it: inStock (true, false, or null when the shop states nothing) beside schema.org's availability (InStock, OutOfStock, PreOrder and the like). The rows say whether a product can be bought, not how many units the shop holds: no inventory count is read.
  • One row per product by default, at its lowest price and in stock when any variant is; switch on one row per variant when you track sizes or colours.
  • Pay only for answers: a page that is blocked, disallowed or carries no product data costs nothing, and so do the requests that list or search a shop; a search that a Shopify or WooCommerce shop answers with no match is an answer, charged like a gone page (see the cost section below).

How to monitor competitor prices on Shopify, WooCommerce and other shops

  1. Open the Actor in Apify Console and press Try for free.
  2. Paste your product page URLs into Product page URLs, one per line — or a shop's home page (or one of its collection or category pages) into Whole shops. The prefilled page shows the shape of a run.
  3. For a whole shop, set Maximum products per shop; switch on One row per variant if you track sizes or colours.
  4. To look for particular products instead of reading the shops whole, list them in Search terms, one per line: a name ("wireless headphones"), a SKU ("SW-10001") or a barcode. Each shop in Whole shops is searched for each term, Maximum products per shop then caps each term's rows in each shop, and every row names its term in searchTerm.
  5. Leave Proxy as it is for a first run, and set the run's maximum total charge if you want a hard cap on cost.
  6. Press Start. The log shows each request, and the run ends with a one-line summary in its status message.
  7. Open the Output tab: the Products view shows one line per product with its image. Export the rows as JSON, CSV, Excel, XML or HTML.
  8. For price monitoring, save the input as a task and schedule it daily or hourly; compare each run's rows with the last to see what moved: a price, a sale's regularPrice, or inStock when a product goes out of stock or comes back.
  9. To let an AI agent use it, add it to Claude, Claude Code, Cursor or any MCP client as a single-purpose connector: https://mcp.apify.com/?telemetry-enabled=false&tools=victorix/product-price-scraper. It loads just this one Actor as a direct tool, so your agent skips the Actor search and uses fewer tokens; new Apify accounts include $5 of monthly usage.

How much does it cost to scrape product prices?

Pay per event: one result event per ok row (one product, or one variant when you ask for variants) and one per not-found row (a page the shop says is gone: HTTP 404 or 410, or a product page that now redirects to the shop's home page; or a search term that no product of a Shopify or WooCommerce shop matched). Rows with invalid-input, source-error or limit-reached are free — a page that is blocked, disallowed by robots.txt, fails to load or carries no product data costs nothing, and listing a whole shop (its home page, feeds and sitemaps) is free too, as is every request a search makes. Platform usage is included in the event price; the one charge that is not per row is a run fee on the apify-actor-start event, $0.000081 per run start at Bronze ($0.081 per 1,000 starts), which Apify charges once up to 1 GB of run memory and once more for each extra gigabyte: a run gets 512 MB for up to 20 pages (its product page URLs plus maxProductsPerShop for each whole shop) and 1024 MB above, the Actor's maximum, so the event is charged once per run. A run you resurrect after it stopped pays one more run start when it starts again (measured on the platform: two apify-actor-start events on each resurrected run).

EventCharged whenPrice (Bronze)Per 1,000
resulta row has status ok or not-found$0.00135$1.35
residential-proxy-per-producta charged row (ok or not-found) was read through a residential IP (below)$0.0009$0.90
apify-actor-starta run starts, or starts again when you resurrect it (once per start, as above)$0.000081$0.081

Rows read through a residential IP. When Apify's datacenter proxy is refused a product page plainly — a 403 or a 429 that shows no sign of a bot challenge, or a dropped connection — the Actor reads the page once more through one of Apify's residential IPs, with the same browser headers as every other request, and that row costs one residential-proxy-per-product event beside its result. A page that shows a challenge (Cloudflare, DataDome, AWS WAF and similar, in its status, headers or content) is never retried that way and stays a free row, and so does a page the residential retry is refused too. With your own proxy URLs in Proxy, or no proxy, the Actor uses no residential IP of its own and the event never applies. A shop whose pages read through residential IPs weigh more than their rows pay for gets no more residential retries in that run (its next refusals are free rows), and the run's status message says how many rows were read through a residential IP.

Worked example, at Bronze: 1,000 product pages, 900 with a product, 30 gone and 70 refused or without product data → 930 result events and one run start, $1.26; if 50 of those 930 rows were read through a residential IP, 50 residential-proxy-per-product events add $0.045. A whole Shopify shop of 1,000 products → 1,000 result events and one run start, $1.35.

Searches. Each product a search term matches is an ok row, charged one result; a product two of your terms match is a row for each. A term that no product of a Shopify or WooCommerce shop matches is one not-found row, charged one result: the shop's catalogue or search answered, and nothing in it matched. That holds when the search stopped at its caps (see Supported shop platforms and limits) and for a barcode too: rows read through Shopify's catalogue (products.json) and WooCommerce's Store API carry no GTIN, so there a barcode finds a product only when the shop uses it as the SKU, while on Shopify's search route each product's own page is read, and a barcode matches when that page states it. A term that matches nothing in any other shop, searched by the addresses of its product pages, is a free source-error row with the code no-match, and a search that could not be finished (a refusal, an error, a page that did not answer) is a free row too. Example: two terms in three Shopify shops, with four matching products in all and two shop-and-term pairs without a match → six result events and one run start.

The table quotes the Bronze price, which is the list price. Silver buyers pay $0.001125 per result ($1.125 per 1,000), $0.00081 per residential row and $0.000072 per run start; Gold buyers $0.0009 per result ($0.90 per 1,000), $0.00072 per residential row and $0.000063 per run start, and so do Platinum and Diamond buyers; the Free tier pays $0.0054 per result ($5.40 per 1,000), $0.0027 per residential row and $0.00009 per run start. Apify's pricing page puts the Free plan on the Free tier, Starter on Bronze, Scale on Silver and Business on Gold; Platinum and Diamond are Apify's enterprise tiers. The free plan's $5 of monthly platform credit covers about 925 results.

The Actor never reads more than your maximum total charge can pay for: it plans only as many pages as the charge limit covers, and when the limit is reached it stops and writes one limit-reached row naming the entry it stopped at and how many are left. A product page whose row lacks a name, a price, a currency or the stock state is read a second time before its row is charged, and the more complete of the two reads is kept; a row whose shop still leaves a field empty (no price, no brand) is an answer and is charged, and the null in that field says what is missing. Rows from a shop's catalogue feed are read once. A page that looks like a block (a challenge page, an empty body) is retried on a fresh session first and becomes a free source-error if it never answers. If a run's compute and residential traffic outgrow what its rows pay for — a very slow or refusing shop, say — the Actor ends early with a free limit-reached row, and everything already delivered stays yours.

How does it compare with other ways to monitor prices?

OptionWhich shopsWhat you getHow you paySchedulingLimits
This Actorindependent shops on Shopify, WooCommerce, Magento, Shopware, PrestaShop, BigCommerce, Wix and Squarespace, from product pages, whole shops or search termsone row per product with the same fields on every platform: price in minor units and as a decimal, ISO currency, regular price, stock, SKU, GTIN, MPN, brand and moreper result, $1.35 per 1,000 at Bronze plus $0.000081 a run, and $0.90 per 1,000 rows read through a residential IP; failed pages free; no monthly minimumApify tasks and schedules, the API, MCP, n8n, Make, Zapierno marketplaces; bot walls and JavaScript-only storefronts end as free rows; no history kept between runs
A marketplace scraperthe marketplaces and large retailers it is built for, such as Amazon or eBaythose sites' own fieldsas its developer prices itdepends on the toolindependent shops only where the tool covers them
A SaaS price monitorthe sites its vendor crawlsdashboards and alerts, often product matching and repricinga monthly plan: entry plans about $49–99 a month, $0.49–0.99 per product a monthbuilt in: one to eight checks a day on the plans reada product cap per plan
Your own scraper per shopthe shops you write it forwhatever you codeyour time, servers and proxiesyours to runone scraper per shop to maintain; it can break when a shop changes its layout
A shop's own product feedone shop at a time: Shopify's catalogue JSON, WooCommerce's Store APIeach platform's own format; neither states a GTINfreeyours to buildone format per platform to parse

The SaaS figures are the entry plans on the vendors' own pricing pages, read on 2026-09-24; this Actor's are the Bronze prices above.

Input

FieldTypeDefaultExample
productUrlsarray of strings (up to 10,000)—["https://shop.example/products/kettle", "shop.example/p/2"]
shopUrlsarray of strings (up to 1,000)—["https://shop.example/"]
searchTermsarray of strings (up to 1,000)—["wireless headphones", "SW-10001"]
maxProductsPerShopinteger 1–100,0001000200
oneRowPerVariantbooleanfalsetrue
includeDescriptionbooleantruefalse
proxyConfigurationApify proxy settings{ "useApifyProxy": true } (datacenter){ "useApifyProxy": false, "proxyUrls": ["http://user:pass@proxy.example:8000"] }

Give at least one product page or shop. A URL typed without https:// is read as https. An entry that is not an http(s) URL, or that repeats an earlier page or shop, gets a free invalid-input row and is never requested. A shop is read from the page you give (its home page is enough), up to maxProductsPerShop products; a collection or category page instead of the home page reads only that listing (see Supported shop platforms and limits); a Shopify collection whose JSON lists no products while the collection counts some is read through the product pages its own page links. Each shop is read once a run, from its first entry: a second entry for the same shop, a second category included, is a free invalid-input row, so run each category of one shop as its own task; product page URLs are not counted against that maximum. Apify's residential proxy group is refused before any request: for residential IPs, give your own proxy URLs. The Actor validates the input before anything is charged; an invalid run writes invalid-input rows and fails without charging.

With searchTerms, every shop in shopUrls is searched for every term instead of being read whole, and product page URLs are read as usual; search terms without a shop to search are refused as invalid-input. A product matches a term when every word of the term is a whole word of the product's name (with the variant's name on a variant row), when the term is the product's SKU or MPN, ignoring case, or when the term is the product's GTIN, its barcode number (8, 12, 13 or 14 digits; spaces and dashes between them ignored). Words are compared without case or accents, so "Café" finds "CAFE", and only whole words count, so "shoe" does not find "shoes". Each shop gives at most maxProductsPerShop rows per term. A term with no letter or digit, or one that repeats an earlier term (ignoring case and spaces), gets one free invalid-input row per shop, which names the term. "How whole shops, categories and searches are read", under Supported shop platforms and limits, says how each kind of shop is searched.

Output: one row per product

One row per product, in the order of your lists: the product pages first, then each shop's products. A page or shop without a product gives one row saying why.

{
"input": "https://warehouse-theme-metal.myshopify.com/products/jbl-charge-3-portable-bluetooth-speaker",
"searchTerm": null,
"status": "ok",
"source": {
"url": "https://warehouse-theme-metal.myshopify.com/products/jbl-charge-3-portable-bluetooth-speaker",
"checkedAt": "2026-09-24T12:00:00.000Z"
},
"error": null,
"result": {
"name": "JBL Charge 3 Portable Bluetooth Speaker",
"variant": null,
"variantCount": 6,
"price": 9995,
"priceDecimal": "99.95",
"regularPrice": null,
"regularPriceDecimal": null,
"currency": "USD",
"availability": "InStock",
"inStock": true,
"sku": "JBL-859042-CHA-BL",
"gtin": null,
"mpn": "50036330442",
"brand": "JBL",
"category": "Portable Speakers",
"imageUrl": "https://warehouse-theme-metal.myshopify.com/cdn/shop/products/30667e89f98c50c9a57a8f10d34d3cc3.jpg?v=1559655757&width=1024",
"url": "https://warehouse-theme-metal.myshopify.com/products/jbl-charge-3-portable-bluetooth-speaker",
"ratingValue": 3,
"reviewCount": 1,
"attributes": null,
"shipping": null,
"shop": "warehouse-theme-metal.myshopify.com",
"shopName": "Warehouse - Metal",
"description": "JBL Charge 3 is the ultimate, high-powered portable Bluetooth speaker with powerful stereo sound and a power bank all in one package. …",
"dataSource": "json-ld",
"scrapedAt": "2026-09-24T12:00:00.000Z"
}
}

The sample's description is cut after its first sentence here; the row carries the whole text, its paragraphs and list lines separated by line breaks. This page states no attributes and no shipping, so both are null.

FieldTypeMeaning
inputstringThe product page or shop URL as you entered it
searchTermstring or nullThe search term the row answers, as you entered it; null on the rows of product pages and of shops read whole
statusok · not-found · invalid-input · source-error · limit-reachedOnly ok and not-found are charged (see the cost section)
source{ url, checkedAt }The page or feed that answered, after redirects (null when none was requested), and when it was read
errornull or { code, message }Why there is no product: no-data, robots-disallowed, blocked, rate-limit, network, http-<status>, unexpected-body, invalid-input, limit-reached and a few more. A no-data message ends with the title of the page that answered (The page's title: “…”., up to 120 characters), so you see what the shop served; a page without product data whose title names a bot check (a CAPTCHA, "access denied", "Pardon Our Interruption") is a blocked row instead. A limit-reached row, always free, says the run stopped early and names the entry it stopped at and how many are left: the run's maximum total charge was reached, its compute outgrew what its rows pay for, or it came near its time limit
resultobject or nullOne product (or one variant), for ok rows; every key on every row

The result fields: name; variant (the variant's own name on a one-row-per-variant row); variantCount; price and priceDecimal; regularPrice and regularPriceDecimal (the struck-through price of a sale, when higher than the price); currency (ISO 4217); availability (one of schema.org's terms such as InStock, OutOfStock, PreOrder); inStock (true, false, or null when the shop states nothing); sku, gtin, mpn, brand; category (the shop's own category, a path joined with " > "); imageUrl (a link, never the picture); url (tracking parameters such as utm_*, gclid and fbclid removed); ratingValue and reviewCount (the shop's aggregate); attributes (colour, size, material and the like, as the shop names them); shipping (the shipping options the page states: country, price, priceDecimal, currency, minDays, maxDays); shop (its host); shopName (the shop's own name); description (plain text); dataSource (json-ld, microdata, rdfa, opengraph, shopify-json or woocommerce-store-api); scrapedAt. A field the shop does not state is null.

How a product becomes one row

  • One row per product. A product with variants is one row at its lowest price (in the main currency), with that variant's regular price; inStock is true when any variant is; variantCount says how many variants it has. With oneRowPerVariant each variant is its own row, linked to its variant (?variant= on Shopify).
  • A product the page describes twice is one row, charged once. When the page's markup states the same product twice (the same name, page and currency, a SKU or GTIN in common, no identifier apart) at two prices, such as without and with VAT, the row carries the price the page's Open Graph tags give the product (product:price:amount, the price with tax on a product page of the PrestaShop shop that printed its product twice, archived in April 2026) when it matches one of the two, else the price stated first; either way with that price's own regular price, and one choice for all of a product's variants.
  • The selling price by its type. Where a page lists a sale price and a struck-through price, the struck-through one (StrikethroughPrice, ListPrice, MSRP) becomes regularPrice, whatever order the page lists them in.
  • A price without an ISO currency is null, because the currency decides how many minor units a decimal is worth.
  • Missing fields are filled from the page's other sources, never guessed: Open Graph product tags fill a markup row, Shopify's own product event fills its stock, and a WooCommerce product page with no product data is answered from the shop's Store API.
  • Products the page points at are not rows: accessories, "often bought with" and related products are what the page links to, not what it sells.
  • The category is the one the product's markup states; else the path of the page's breadcrumbs, without the home page and the product itself ("Audio > Speakers"); else the page's Open Graph product:category. Shopify's catalogue gives the product type, WooCommerce's Store API the product's categories, several joined with ", ".
  • Attributes as the shop names them: the markup's additionalProperty entries and the product's color, material, size and pattern; Shopify's options; WooCommerce's attributes. A variant row carries its own values, a product row every variant's values joined with ", " ("Color": "Blue, Red, Black"). A product whose shop states none has attributes: null, never an empty object.
  • Shipping as the page states it (schema.org OfferShippingDetails): one entry per shipping option, and per destination country when an option names several (country is the two-letter code as the page states it, EU included, or null when the option names none), the rate in the currency's minor units (0 is free shipping), and minDays/maxDays as handling time plus transit time. A page that states no shipping gives shipping: null, and so do all rows read through Shopify's and WooCommerce's feeds, which carry none.
  • The shop's name on a row read from a page is its og:site_name, else the name of its JSON-LD Organization or WebSite, else the offer's seller when that is a business, never a person; a row read through Shopify's catalogue (a whole shop, a collection, or a search through the catalogue) carries the name the shop's /meta.json states, and a row read through WooCommerce's Store API carries none, because the API names no shop.
  • The description as plain text: tags removed, entities decoded, paragraphs and list items on lines of their own, at most one blank line between paragraphs, no length cap. It comes from the product's markup, Shopify's body_html or WooCommerce's description (else its short description), and last from the page's og:description. Turn includeDescription off for smaller rows: every row's description is then null, and each row is charged the same.

Supported shop platforms and limits

PlatformProduct pagesWhole shopOne collection or categorySearch terms
Shopifyschema.org markup and Open Graph tags; stock from Shopify's own product event when the markup lacks itthe catalogue JSON, 250 products a requestthe collection's JSONthe catalogue, or the shop's search page and each product's page
WooCommerceschema.org markup and Open Graph tags, else the shop's Store APIthe Store API, 100 products a requestthe Store API's category filterthe Store API's search and SKU filter
Magento, Shopware, PrestaShop, BigCommerce, Wix, Squarespaceschema.org markup and Open Graph tagsthe product pages their sitemaps listthe product pages their sitemaps list under the category's addressthe addresses of those product pages (best effort)
  • Measured coverage. On 154 product pages of 46 independent shops across eight platforms (Shopify, WooCommerce, Wix, Squarespace, BigCommerce, PrestaShop, Magento, Shopware), read live on 2026-09-24, 87.6 % answered a complete row (name, price, ISO currency and stock; 134 of the 153 pages still online). The rest: bot walls (10 pages of three shops), a headless storefront that publishes no product data (4), two pages a shop served without its product data at the time, two rows whose shop states no stock, and one site whose robots.txt closes it to all robots. Of 19 whole shops, 14 listed their products; the others were two bot walls, the headless storefront and two shops whose servers answered web pages at their sitemap addresses during the run.
  • Bot walls are not passed. A shop that answers with a challenge or a captcha (Cloudflare, AWS WAF, SiteGround, DataDome and similar) gets a free source-error row, whether the page says so in its body or in the status and headers the vendor documents. The Actor never solves challenges, and it stops asking a shop that refuses most of its latest requests. A plain refusal on Apify's datacenter proxy (a 403 or a 429 with no sign of a challenge, or a dropped connection) gets one more try through a residential IP (see the cost section); a challenge never does.
  • Only what the server sends. The Actor reads the HTML and JSON a shop's server returns without running JavaScript: a storefront that writes its product data with JavaScript after the page loads (some headless shops) answers no-data, free.
  • Big marketplaces are out of scope. It is built for independent online shops; Amazon, eBay and the large retailers mostly refuse polite automated requests.
  • The shop's own statement. Rows report what the shop publishes; a price shown on the page but not in its data is not read. A shop that sells in several currencies answers in the market it serves to the proxy's location.
  • Politeness: at most 120 requests a minute for the whole run, one second between two requests to the same site, two retries per page, and one more through a residential IP after a plain refusal (see the cost section); every site's robots.txt is read first and obeyed.
  • Run time: about a second per product page from one shop, and up to 120 pages a minute across many shops (two a second on average), read side by side, because each site gets one request a second at most and the whole run 120 a minute; a product page whose data is partial is read twice, so it takes about two seconds, or about four when a WooCommerce page also asks the shop's Store API on each read. A whole Shopify shop of 1,000 products takes about ten seconds (four catalogue requests and a few to find them); a shop read through its sitemap takes about a second per product.
  • Time limit. A run has this Actor's default time limit of one hour (3,600 seconds) unless you set another in the run options. Shortly before it — 60 seconds before, or a fifth of the time limit when that is shorter — the Actor stops taking new pages, delivers and charges what it has read, and writes one free limit-reached row naming the entry it stopped at and how many are left; the run ends SUCCEEDED (unless it had answered nothing yet and a page failed, which ends a first run FAILED), and resurrecting it carries on from that entry. In an hour a run reads about 7,000 product pages (120 requests a minute); Shopify and WooCommerce catalogues answer 250 or 100 products a request, so 300 such shops of 1,000 products each need about 3,000 to 4,000 requests, 25 to 35 minutes. For more, raise the time limit or split the list across runs.
  • Freshness: every row carries the minute it was read; nothing is cached across runs. Rows go out in batches, so a run that is aborted or times out keeps every finished batch.
  • Memory: a run of up to 20 pages (its product page URLs plus maxProductsPerShop for each whole shop) gets 512 MB and a bigger run 1024 MB, worked out from your input just before the run starts; you can still set the memory yourself in the run options, and long lists of product pages and shops read through their sitemaps finish sooner at 1024 MB. The Actor keeps Node's heap within three quarters of the run's memory, so its garbage is collected before the run outgrows its memory. Either size costs you the same, per result plus the one start event: measured on the platform on 2026-09-25, 154 product pages from 46 shops took 178 s at 1024 MB against 304 s at 512 MB, and two whole shops read through their sitemaps, 100 products, took 181 s against 289 s.

How whole shops, categories and searches are read

  • Whole shops: Shopify shops are read through their catalogue (all published products), WooCommerce shops through their Store API, other shops through at most 25 sitemap files. A sitemap named for products is read whole; in a sitemap that mixes products with categories and content pages, a product is known by its address (PrestaShop's /123-name.html, /products/…, /p/…), else by what the platform marks its products with (Magento and PrestaShop list a product's images, Shopware marks products hourly). Only the language the shop's home page answers in is read: the same products under /fr/ or in a shop's French sitemap are the same rows again. A path prefix counts as another language when the home page declares it in its hreflang links, or, when it declares none, when the home page itself answers under a language prefix such as /en/; so at a shop that declares its languages or answers at the root, a category such as /it/ or /hi-fi/ stays a category. The rare shop that answers at the root, declares no languages and still lists its other languages' pages under prefixes has those pages read too, each a row of its own. A shop that refuses its sitemap to us gets a free blocked row, a shop whose sitemap files answer with something that is not a sitemap a free unexpected-body row, and a shop without one a free no-data row: list its product pages instead. A sitemap file may take up to 80 seconds to arrive (a page 30); one still arriving then is retried, and the run log names each sitemap file read with its size and the product pages it gave.
  • One collection or category: give its page instead of the home page and only that listing is read, cut at maxProductsPerShop and charged per product row like a whole shop. A Shopify collection (/collections/{handle}, a language prefix such as /fr/ allowed) is read through that collection's JSON, 250 products a request. When that JSON lists no products but the collection's own record (/collections/{handle}.json) counts some, the product pages the collection's page links (/products/{handle} on the same shop: the theme's product cards when it marks them; links outside the page's main content or inside its header, nav, aside or footer elements are left out; one per product, only what that one page shows) are read instead, up to maxProductsPerShop, each checked against robots.txt and charged like any product page; /collections/all is the whole shop, as are a WooCommerce shop's shop page (/shop/) and the front page of a WordPress shop installed under a folder (/store/). A WooCommerce product category (the page WooCommerce prints for it, whatever its address) is read through the Store API's category filter, 100 a request. Any other shop's category page is read through the product pages its sitemaps list under that page's path: /shop/audio/ reads /shop/audio/…, and a category named like a file, Magento's /audio.html, reads /audio/…. A shop whose product addresses do not sit under the category's (PrestaShop's /3-clothes lists /men/1-1-t-shirt.html), a WooCommerce tag, brand or attribute page, and a collection whose JSON is closed to us each get one free no-data row saying the category could not be listed (a blocked row when the shop also refuses its sitemap to us), never the whole shop; so does a collection or category that lists no products (a Shopify collection whose JSON lists none and whose record counts none, cannot be read, or whose page links no product page), a category page that now redirects to the home page or the all-products page gets a free no-data row (the category is usually gone), and one that answers 404 a free http-404 row.
  • Search in a Shopify shop: a shop whose catalogue takes at most two requests of 250 products per term (by the count its /meta.json states) is read whole once, for all of your terms, and every product and variant in it is matched; when the last page that count plans comes back full, the next is read too, while pages stay full, so a count that is out of date leaves nothing unread. The catalogue states no GTIN. A bigger shop, or one that states no count, is searched through its own search page (/search?q=…&type=product, out-of-stock products listed last where the merchant allows it), up to 10 pages per term, and each product a page links is read through its product page (/products/{handle}), as a product page you list is, up to maxProductsPerShop products per term: the page states each variant's stock and currency, and often its GTIN, so a barcode matches there when the page states it. When no term's first search page links a product, or the shop's robots.txt closes its search page, the shop is read through its catalogue instead. A collection page with search terms is read through that collection's JSON, which the search page cannot be narrowed to. Shopify's own search can leave out products the merchant hides from it, a shop that uses a search app answers as that app does, and near matches it lists are dropped by the strict rule under Input.
  • Search in a WooCommerce shop: the Store API is searched for the longest word of each term, in lower case and without accents, and, for a term without spaces, asked for it as a SKU too; the products it returns, up to 10 pages of 100 and at most maxProductsPerShop per term, are matched, and a product category page searches inside that category. A shop whose database compares accents strictly may not return a product whose name has them.
  • Search in any other shop (best effort): the product pages the shop's sitemaps list are taken as a whole-shop read finds them, up to 100,000 whatever maxProductsPerShop says (a product may be anywhere in them), and for each term only the pages whose address holds every word of the term, or the SKU (SW-10001 finds /…/SW10001), are read, up to maxProductsPerShop pages per term, each matched on what it states. A product whose page address does not name it is missed, which is why a term that matches nothing there is a free no-match row; a page the shop says is gone, and one that could not be read, are counted in that row's message. A Shopify or WooCommerce shop whose catalogue does not answer or lists nothing is searched this way too, and so is a Shopify collection read through the product pages its own page links.
  • Variants in a search: each variant is matched on its own, in Shopify's catalogue and on every product page a search reads, so a term that names one size, colour, SKU or barcode finds that variant. With one row per product, the product's row covers only the variants that match: its lowest price among them, in stock when one of them is, variantCount counting them, and the SKU, GTIN and MPN they share; a single matching variant is that variant's own row.
  • A search's caps: per shop and term, at most 10 pages of search results and maxProductsPerShop products read (in a best-effort search, maxProductsPerShop product pages), and at most maxProductsPerShop rows; a Shopify catalogue read whole for a search stops after 400 requests (100,000 products). A common word in a big shop can list more products than that, and a match beyond them is missed; a term whose search stopped at its caps without a match is still a charged not-found in a Shopify or WooCommerce shop (see the cost section). A run takes up to 1,000 search terms, so that the progress it saves to carry on after the platform moves it to another server stays small; split a longer list across runs.

The Actor reads only public product facts that shops publish for search engines, obeys every site's robots.txt, paces its requests, and never logs in or solves a challenge; whether your use of a site's data is allowed by that site's terms is yours to check.

The data comes from the pages and shops you list, and only from the product data they publish: schema.org markup and Open Graph tags on the page, Shopify's catalogue JSON (which Shopify documents for read-only agents in a store's /agents.md), WooCommerce's public Store API, and the sitemaps a shop's robots.txt names or keeps at the usual addresses. A search also reads a Shopify shop's own search page (/search?q=…&type=product, which /agents.md lists for read-only agents too) and the product page of each product it lists (/products/{handle}), and the Store API's search and sku filters, which WooCommerce documents; in any other shop it reads only the product pages its sitemaps list. Only product data is read — names, prices, stock, identifiers, brands, categories, attributes, shipping options, ratings, the shop's name and the product's description, which is delivered as plain text unless you switch includeDescription off — never reviews, reviewer names, images or personal data. The Actor:

  • respects robots.txt: it reads each site's robots.txt before the first page, applying the rules for all robots and any group for its technical name, and never requests a disallowed page, whose row says so (robots-disallowed, free); a Shopify shop whose robots.txt closes its search page is searched through its catalogue instead; when Apify's datacenter proxy was refused a site's robots.txt plainly (a 403 or a 429 with no challenge), the file is read through a residential IP before any page of that site is read that way, and a site that answered it with a challenge or a login gets no page read that way;
  • reads only public pages: no login, no cookies of yours, and no CAPTCHA or bot-challenge solving — a blocked page is a free source-error row, and a page that shows a challenge is never retried through a residential IP;
  • paces its requests (see Supported shop platforms and limits), and the residential retry carries no cookie.

You decide which pages and shops it reads: make sure your use of each site's data is allowed by that site's terms. Not affiliated with, sponsored or endorsed by Shopify, WooCommerce, WordPress, Wix, Squarespace, BigCommerce, PrestaShop, Adobe (Magento), Shopware, schema.org, Google, Meta or any shop you read with it. The vocabulary is schema.org's (https://schema.org/Product); prices are converted with the ISO 4217 minor units.

Support: open an Issue on the Actor's Issues tab and you get a reply within 12 hours.

Integrations: API, MCP, n8n, Make, Zapier and Google Sheets

  • API: start a run with POST https://api.apify.com/v2/acts/victorix~product-price-scraper/runs, the JSON input as the body and your token in the Authorization: Bearer <token> header, never in the URL; then read the run's default dataset as JSON or CSV with the same header. The Actor's API tab in Console shows each call ready to copy.
  • AI agents: the Actor is available to AI agents through Apify's MCP server, which can call any public Actor by name; add it as a connector with https://mcp.apify.com/?telemetry-enabled=false&tools=victorix/product-price-scraper. Each tool call is one run, charged one result event per product.
  • n8n, Make, Zapier and Clay: use the Apify node, module or integration, pick this Actor and pass productUrls or shopUrls as arrays; send the rows to Google Sheets, Slack or your database.
  • Price-drop alerts from a schedule: save your input as a task, schedule it, and compare each run's price and inStock with the last. The rows carry a URL, so a task's last run can also be read as an RSS feed from Console — treat that feed link like a password, because it carries a token.

More Actors from us

Every public Actor of ours is listed on our Apify profile: apify.com/victorix.

FAQ

Which shop platforms does it support?

Independent shops on eight platforms: Shopify, WooCommerce, Magento, Shopware, PrestaShop, BigCommerce, Wix and Squarespace, all measured live on 2026-09-24. Product pages on all eight are read from their markup; whole Shopify and WooCommerce shops through their own catalogue feeds, the other six through the product pages their sitemaps list. A product page on another platform is read the same way when it carries schema.org Product markup or Open Graph product tags.

How do I monitor competitor prices every day?

Put the competitor product pages (or their whole shops) into the input, save it as a task and schedule it daily in Apify Console. Each run writes one row per product with price, regularPrice and inStock; compare a run's rows with the previous run's, or send them to Google Sheets through n8n, Make or Zapier and let a formula flag what moved.

Can it tell me when a product is back in stock?

Yes, by comparing runs. Every row carries inStock and schema.org's availability, so schedule a task and compare each product's inStock with the previous run's: false to true is back in stock, true to false is out of stock. The Actor itself keeps no history between runs, so the comparison happens in your sheet, database or n8n, Make or Zapier flow.

Can it scrape a whole Shopify or WooCommerce shop?

Yes. Put the shop's home page into Whole shops. Shopify shops are read through their catalogue JSON, 250 products a request; WooCommerce shops through their Store API, 100 a request; any other shop — Magento, Shopware, PrestaShop, BigCommerce, Wix, Squarespace — through the product pages its sitemaps list. maxProductsPerShop caps each shop, and listing the shop is free: you pay per product row.

Can it read one category or collection of a shop?

Yes. Put the collection or category page into Whole shops instead of the home page: only that listing's products are rows, up to maxProductsPerShop, each charged like a whole shop's, and a category the Actor cannot list is one free row, never the whole shop. Each shop is read once a run, so to track two categories of one shop, save each as its own task.

Can it find a product by name, SKU or barcode?

Yes. Put the shops into Whole shops and what you look for into Search terms, one per line: each shop is searched for each term instead of read whole, and only the matching products become rows, each naming its term in searchTerm. A product matches when every word of the term is a whole word of its name, or when the term is its SKU, MPN or barcode (GTIN).

Does it work with EAN or UPC barcodes?

Yes: EAN-13 and UPC-A are both GTINs, and gtin carries the one a product page states. A barcode search term matches when its digits equal the product's GTIN once both are padded to 14 digits, so a 12-digit UPC finds a product listed under the same number as a 13-digit EAN. Shopify's catalogue and WooCommerce's Store API state no GTIN, so there a barcode matches only a SKU: for barcodes, the name or the SKU is the surer search.

How is it different from marketplace scrapers?

A marketplace scraper is built for big marketplaces and retailers such as Amazon or eBay; this Actor reads independent shops on eight platforms, many shops in one run, and gives every row the same fields whatever the platform. It does not read marketplaces: their pages usually end as free source-error rows, so for a marketplace, use a scraper built for it.

Does it work on Amazon or eBay?

It is built for independent online shops. Amazon, eBay and most large retailers refuse polite automated requests or hide their data behind bot walls, which this Actor does not pass; their pages usually end as free source-error rows.

How long does a run take?

About a second per product page from one shop, and up to 120 pages a minute across many shops read side by side, because each site gets at most one request a second and the whole run 120 a minute. A whole Shopify shop of 1,000 products takes about ten seconds; a shop read through its sitemap about a second per product.

What happens when a run reaches its time limit?

It stops cleanly just before it: 60 seconds before the limit, or a fifth of the time limit when that is shorter, it stops taking new pages, delivers and charges what it has read, writes one free limit-reached row naming the entry it stopped at and how many are left, and ends SUCCEEDED, unless it had answered nothing yet and a page failed. Resurrect the run to carry on from that entry, or raise the time limit.

Why was a search that found nothing charged?

In a Shopify or WooCommerce shop the catalogue or search answered and nothing in it matched, which is an answer: a not-found row, charged like a gone page, also when the search stopped at its caps or the shop did not state the barcode where the Actor read it. In any other shop a term that matches nothing is a free no-match row, and a search the shop refused, or that could not be finished, is always free.

Why did a search miss a product the shop sells?

The match is strict: every word of the term must be a whole word of the product's name ("shoe" does not find "Running Shoes"), or the term must be its whole SKU, MPN or barcode. A Shopify shop's search also leaves out products its merchant hides, a common word in a big shop can list more products than a search reads, and a shop searched by page address never reads a product whose address does not name the term.

Why did a page answer no-data?

The page answered, but it carries no product data this Actor reads: no schema.org Product markup and no Open Graph product tags, often because the shop writes them with JavaScript. The row is free, and its message ends with the title of the page that answered: a shop's own product name there means the page came without its data (read it again later), an unexpected title means the shop served something else.

Why is a price or stock value null?

The shop's data did not state it, or stated a price without an ISO 4217 currency code, which cannot become minor units without guessing. The Actor fills a missing field from the page's other sources and reads such a page twice before charging it. A WooCommerce product the shop shows but does not sell, with no price printed, has a null price too: WooCommerce reports it as 0, which is not what it costs.

Why did a row cost a residential-proxy-per-product event?

Apify's datacenter proxy was refused that product page plainly — a 403 or a 429 with no sign of a bot challenge, or a dropped connection — so the Actor read it once more through one of Apify's residential IPs, which cost more than a datacenter request. A challenge is never retried that way. To never pay the event, choose your own proxies, or none, in Proxy.

Changelog

  • 0.2 (2026-09-29): a long run keeps its memory flat: each page's request now leaves memory once it has been read, where every one used to be kept to the end of the run, about 0.8 KB a page (100,000 pages measured locally: the memory kept stayed level instead of growing by about 70 MB, and the local run took under 5 minutes instead of 27). A product page whose one residential retry was still due when a run stopped near its time limit is left for the resurrected run to read instead of ending as a free refusal row. A long list at a forced 512 MB could end after its first rows with "nothing was charged", because its first rows were read but not yet charged when the cost guard first looked; the guard now counts rows read and not yet charged.
  • 0.2 (2026-09-28): the Store title is now Product Price & Stock Scraper: Shopify, WooCommerce (was Product Price & Stock Scraper for Online Shops). Node's heap is kept within three quarters of the run's memory: the base image allowed it 30 GB, so a run of a few thousand pages could outgrow a 1 GB run and stop without an error. A run near its time limit stops cleanly: it delivers and charges what it has read, writes one free limit-reached row naming the entry it stopped at, and ends SUCCEEDED instead of TIMED-OUT.
  • 0.2 (2026-09-24): renamed Product Price & Stock Scraper (was Schema.org Product Extractor). Whole shops through Shopify's catalogue JSON, WooCommerce's Store API and sitemaps (products found in the mixed sitemaps of Magento, Shopware and PrestaShop too, in the shop's own language); one row per product by default, one per variant on request; 21 result fields (decimal prices, regular price, stock flag, GTIN, MPN, image, rating, shop, data source); Open Graph tags and Shopify's product event read; a page without product data is now free; challenges that announce themselves in their status or headers are free blocked rows; robots.txt read by the Actor itself; a margin guard and a per-shop block-rate stop; one Shopify collection, WooCommerce product category or other shop's category page read on its own, from its page in Whole shops; a no-data row names the title of the page that answered, and a page without product data titled like a bot check is a free blocked row; a WooCommerce product the shop shows but does not sell (not purchasable, no printed price) has a null price instead of 0, and like any row with a missing field it is still charged; a product page with a missing name, price, currency or stock state is read a second time before its row is charged, and the more complete read is kept. The Actor paces each shop itself, so a run over many pages of one shop reads about one a second, as Supported shop platforms and limits says, instead of one every few seconds. A Shopify collection whose JSON lists no products while the collection's record counts some is read through the product pages its page links, instead of one free row. A run over several shops reads them side by side instead of one shop after another (three shops of 20 pages: 20 s, was 40 s). The pace holds on the wire too: a site's first page waits a second after its robots.txt answered, a retry a second after the failed attempt, and the Actor spaces its requests and holds its 120 a minute once more just before each connection (a live sample of 58 entries: no two requests to one site under a second apart and at most 120 requests in a minute; the version before measured gaps from 855 ms and, reading many shops side by side, 125 requests in a minute). A product page that Apify's datacenter proxy refuses plainly (a 403 or a 429 with no sign of a bot challenge, or a dropped connection) is read once more through a residential IP, for one residential-proxy-per-product event beside its result; a page that shows a challenge, DataDome's included, never is, and a refused page's row now says which of the two it met. A run's memory now follows its size, 512 MB for up to 20 pages (product page URLs plus maxProductsPerShop for each whole shop) and 1024 MB above, so long lists and shops read through their sitemaps finish sooner (154 product pages: 178 s instead of 304 s) at the same price. A product the page's markup states twice at two prices (such as without and with VAT, with a SKU or GTIN in common) is one row, charged once, at the price its Open Graph tags give (product:price:amount) when that is one of the two, else at the price stated first. Five more result fields: category, attributes, shipping, shopName (on Shopify's catalogue rows too, from the shop's /meta.json) and description, the product's description as plain text, on by default and switched off with includeDescription. Keyword search: searchTerms looks for products by name, SKU or barcode in the shops of shopUrls instead of reading them whole — Shopify shops through their catalogue or their own search page, WooCommerce shops through the Store API's search, other shops through the addresses of their product pages (best effort) — and every row carries the new top-level searchTerm; a term no product of a Shopify or WooCommerce shop matches is a charged not-found, and one that matches nothing in a shop searched by page address a free no-match row. Prices are 90 % of the niche leader's on each of Console's four tiers, and Platinum and Diamond buyers pay the Gold prices: $1.35 per 1,000 results at Bronze, $0.90 per 1,000 rows read through a residential IP and $0.000081 per run start (see the cost section). A run's saved progress now stays small whatever the run's size: each shop's product list is saved on its own as the run lists it, and only the rows about to go out wait in the saved progress, so a run over many large shops or many search terms can move to another server and carry on without sending or charging a row twice, as a small run always could; a run takes up to 1,000 search terms. The progress is also saved after every batch of rows, so a run that stops without warning (its server fails or runs out of memory) and is resurrected sends and charges again at most the batch it was sending, where before it could repeat every row since the last save, made once a minute. A sitemap file that arrived too slowly to finish in 30 seconds was read as far as it had come and skipped without a word, so a large shop's sitemaps could give no product at all: a sitemap file now gets 80 seconds, a page still arriving when its time runs out is retried instead of read, and a shop whose sitemap files cannot be read says so in its row. The run's closing status message counts the result events the platform charged; it had said twice the number, because the SDK's own count adds the dataset-item event it books on every row.
  • 0.1 (2026-09-23): first version. The parser reads JSON-LD, microdata and RDFa; the crawler respects robots.txt and reads only as many pages as the charge limit pays for.