# E-commerce Product Scraper for Any Online Store (`digital_influx/product-data`) Actor

Products from any online store: price, sale price, stock, variants, SKU, GTIN, rating, images. Give a store to read its whole catalog, or product pages. Schedule it to get price drops, stock changes, new and removed products. Shopify, WooCommerce and more.

- **URL**: https://apify.com/digital\_influx/product-data.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** E-commerce, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 product reads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## E-commerce Product Scraper for Any Online Store

Get the products of **any online store**: name, brand, SKU, GTIN, price, sale price and discount, currency, stock, **every variant with its own price and stock**, rating and reviews, images and description.

- **A whole catalog:** give the store (`thesill.com`) and the Actor finds every product page in the store's sitemap.
- **Just some products:** give product page URLs from any store.
- **Price monitoring:** schedule it daily or weekly and each product says what changed since the last run: price drops and increases with the amount and percent, back in stock, sold out, new products, and products removed from the store.

It reads the product data that stores publish for Google and link previews: schema.org JSON-LD, microdata, Open Graph product tags, and the product object every Shopify page carries. No browser, so it is fast and cheap, and no guessing from the page layout: a value is either published by the store or `null`.

### What it works on

Checked live on 2026-09-25 against 31 stores:

| Works | Examples |
|---|---|
| **Shopify** (with or without JSON-LD in the theme) | The Sill, Brooklinen, Stumptown, Cotopaxi, Natori, Tentree, Cosmoparis, Burrow |
| **WooCommerce** | Nalgene |
| **Salesforce Commerce Cloud** stores that print JSON-LD | Christopher Ward |
| BigCommerce, Magento, PrestaShop, Shopware, Wix, Squarespace, Tiendanube, VTEX and custom stores | whenever the product page carries schema.org, Open Graph or microdata product data |

**What it does not do:** big retailers that block automated visits (REI, Decathlon, LEGO, Newegg and others answered 403 in the test) and pages built only by JavaScript. The Actor does not try to get around a block: those pages are reported as `blocked_by_site` or `no_product_data` and **cost nothing**.

### Input

| Field | What it does |
|---|---|
| Store URLs | Whole stores, one per line (`thesill.com`). The Actor finds the product pages in the store's sitemap. |
| Product URLs | Single product pages from any store. |
| Max products per store | 100 by default, up to 5,000. Product pages you list yourself are always read. |
| Compare with the previous run | For schedules: each product gets `changes` against the last time it was read. |
| Only changes | With monitoring on, unchanged products are not saved (they are still read and charged). |
| Minimum price change (%) | Smaller price changes are not reported; the last reported price stays the reference, so small changes that add up are reported once the total passes it. |
| Memory name | Optional: a separate price history, for example one per client, or a new name to start over. |
| Variants / Description | Each size or color with its own price, SKU, GTIN and stock; the first 1,000 characters of the description. Both on by default. |

```json
{
  "storeUrls": ["https://www.thesill.com", "brooklinen.com"],
  "productUrls": ["https://www.cotopaxi.com/products/allpa-50l-duffel"],
  "maxProductsPerStore": 500,
  "compareWithPreviousRun": true,
  "minPriceChangePercent": 5
}
```

### Output

One row per product:

```json
{
  "type": "product",
  "store": "thesill.com",
  "url": "https://www.thesill.com/products/plant-mom-tote-bag",
  "status": "ok",
  "name": "Plant Mom Tote Bag",
  "brand": "The Sill",
  "sku": "AC-TBG-PM-NT",
  "gtin": null,
  "productId": "9400312081",
  "category": "Accessory",
  "price": 15,
  "currency": "USD",
  "listPrice": 19,
  "onSale": true,
  "discountPercent": 21.1,
  "minPrice": null,
  "maxPrice": null,
  "availability": "out_of_stock",
  "inStock": false,
  "rating": 5,
  "reviewCount": 1,
  "variantCount": 1,
  "variants": [
    { "id": "34400340561", "name": "Plant Mom Tote Bag", "sku": "AC-TBG-PM-NT", "gtin": null, "price": 15, "listPrice": 19,
      "currency": "USD", "availability": "out_of_stock", "url": "https://www.thesill.com/products/plant-mom-tote-bag?variant=34400340561" }
  ],
  "image": "https://www.thesill.com/cdn/shop/products/the-sill_plant-mom-tote_variant_01.jpg",
  "platform": "shopify",
  "dataSources": ["json-ld", "shopify", "open-graph"],
  "checkedAt": "2026-09-25T20:32:33.842Z"
}
```

- `price` is the product's price, or the lowest one when variants differ (`minPrice` and `maxPrice` then give the range).
- `listPrice` is the price before a sale (a strikethrough, list or compare-at price), only when it is higher than `price`.
- `availability` is one of `in_stock`, `limited_stock`, `out_of_stock`, `pre_order`, `back_order`, `made_to_order`, `in_store_only`, `discontinued`, `reserved` or `null` (not published). `inStock` is `true` when any variant can be bought.
- `gtin` is only filled when its check digit is valid; some stores put other numbers in that field.
- `dataSources` says where the data came from.

When you give a store, a summary row per store (`type: "store"`) says how many product pages its sitemap lists, how many were read, and `truncated: true` when there were more than the limit.

Rows for pages you listed that could not be read are free and say why: `blocked_by_robots`, `blocked_by_site` (the store answered 401/403), `not_found`, `no_product_data`, `error` or `invalid`.

### Price monitoring: only what changed

Turn on **Compare with the previous run** and schedule the Actor. Each product then has:

```json
"changes": { "status": "changed", "price": { "from": 10, "to": 8, "change": -2, "percent": -20 }, "availability": null, "currency": null }
```

`status` is `new` (the first time this product is checked: when the whole store fits in the limit, a product added to the store), `changed` or `unchanged`. The store row adds `changes` with the counts of new products, price drops, price increases, products back in stock and sold out, and the products **removed** from the store since the last run. Removed products are only reported when the whole catalog was read, so a run cut by the limit never claims a product disappeared.

- **Only save products that changed** keeps the dataset small for alerts: unchanged products are still checked, and charged, but not saved.
- **Minimum price change (%)** ignores small moves. The last reported price stays the reference, so small changes that add up are reported once the total passes the threshold.
- Give a **Memory name** to keep separate histories, for example one per client.
- The first run has nothing to compare with (`changes: null`).
- Some stores show prices in the visitor's currency. If a store switches currency, the change is reported as `currency: { from, to }` and the prices are not compared.

### Pricing

**Pay per event:** one `product` event per product page read with product data. Store summary rows, pages without product data, blocked pages and errors are free. Set a maximum charge per run in Apify and the Actor stops there.

### How it reads a store

- It asks `robots.txt` first, for every store and every page: pages it disallows are not read, and a `Crawl-delay` is respected.
- Pages are read at most 2 at a time per store, and 429 or 503 answers are retried after a pause.
- Product pages come from the store's product sitemap (`sitemap_products_1.xml` on Shopify, `product-sitemap.xml` on WooCommerce, and others). Without one, the Actor tries the sitemap pages that are not blog posts, categories or other obvious non-product pages.

Use it on stores whose terms allow reading their public pages, such as your own store, stores that have given you permission, or public price information you are allowed to collect. The User-Agent names the Actor, so store owners can see it and allow or block it in their robots.txt.

### FAQ

**Can it read Amazon, Walmart or other marketplaces?** No. They block automated visits and their terms forbid it.

**Why is a price `null`?** The store does not publish a price in its product data (for example, "price on request"). The Actor never guesses a price from the page layout.

**Does it follow a category page?** Give the store address instead: the sitemap lists every product, including ones not shown in any category.

# Actor input Schema

## `storeUrls` (type: `array`):

One store per line, for example allbirds.com. The Actor finds the product pages in the store's sitemap (Shopify, WooCommerce, BigCommerce, Salesforce and most platforms publish one) and reads up to the limit below.

## `productUrls` (type: `array`):

One product page per line, from any store. Pages without product data are saved free with a status that says why.

## `maxProductsPerStore` (type: `integer`):

How many products to read from each store (1 to 5000). Product pages you list yourself are always read.

## `compareWithPreviousRun` (type: `boolean`):

For scheduled runs: each product gets "changes" (new product, price change with amount and percent, stock change) against the last time it was read, and each store row counts price drops, increases, products back in stock or sold out, and products removed from the store. What each product cost is kept in a key-value store named "product-data-state" in your account.

## `onlyChanges` (type: `boolean`):

With monitoring on: products whose price and stock did not change are not saved (they are still read and charged, since the check is the work).

## `minPriceChangePercent` (type: `integer`):

With monitoring on: price changes smaller than this percent are not reported (0 = every change). The last reported price stays the reference, so small changes that add up are reported once the total passes the threshold.

## `stateKey` (type: `string`):

Give a name to keep a separate price history, for example one per client, or a new name to start over.

## `includeVariants` (type: `boolean`):

Each size, color or pack with its own price, SKU, GTIN and stock (up to 100 per product).

## `includeDescription` (type: `boolean`):

The product description (first 1000 characters).

## `maxConcurrency` (type: `integer`):

How many stores are read at the same time (1 to 10). Each store is read at most 2 pages at a time, or one page at a time with the pause its robots.txt asks for (Crawl-delay).

## Actor input object example

```json
{
  "storeUrls": [
    "https://www.thesill.com"
  ],
  "maxProductsPerStore": 20,
  "compareWithPreviousRun": false,
  "onlyChanges": false,
  "minPriceChangePercent": 0,
  "includeVariants": true,
  "includeDescription": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "storeUrls": [
        "https://www.thesill.com"
    ],
    "maxProductsPerStore": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/product-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "storeUrls": ["https://www.thesill.com"],
    "maxProductsPerStore": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/product-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "storeUrls": [
    "https://www.thesill.com"
  ],
  "maxProductsPerStore": 20
}' |
apify call digital_influx/product-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/product-data"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sdpACak2EEWFG6RMW/builds/JChTwyxmHTngSASIS/openapi.json
