# Shopify Product Monitor & Scraper - Price, Stock & Sales (`lukehunter/shopify-store-products-scraper`) Actor

Scrape Shopify products or monitor stores for product additions/removals, price changes, sales and stock changes. Stateful scheduled runs emit only change events plus free store summaries. Public product feeds only.

- **URL**: https://apify.com/lukehunter/shopify-store-products-scraper.md
- **Developed by:** [Luke Hunter](https://apify.com/lukehunter) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 products

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Shopify Product Monitor & Scraper — Price, Stock & Sales

**For e-commerce competitor monitoring, dropshippers sourcing products, and market researchers who need clean Shopify catalogue data or compact change alerts.** In normal mode it returns every product (or variant). Turn on `onlyChangesSinceLastRun` for persistent monitoring: the first run saves a baseline, later runs emit only price, sale, stock and assortment changes. No login, no API key, no browser.

Pay-per-result: **$0.002 per delivered product — 1,000 products = $2.**

### Quick start (1 minute)

1. Open the **Input** tab and list the store(s) you want, e.g. `https://www.allbirds.com`.
2. Click **Start**. Export the resulting dataset to CSV/Excel/JSON, or pull it through the API shown below.

```json
{
  "storeUrls": ["https://www.allbirds.com"],
  "maxItems": 20
}
```

That prefilled example finishes in well under a minute and costs $0.04.

### How it works

Every Shopify store publishes a public, unauthenticated JSON feed of its catalogue at `/products.json` (and, per collection, at `/collections/<handle>/products.json`). This Actor:

1. Checks the store's `robots.txt` for the exact path it is about to read, and skips the store (reporting why, in the run's status message) if that path is disallowed.
2. Pages through `products.json?limit=250&page=N`, stopping on the first empty page — the real end of the catalogue.
3. Normalises every product (or, in variant mode, every variant) into one flat dataset row.
4. Reads store domains politely: about one request per second per store, backing off and retrying briefly on rate limiting, and reporting — never disguising — a block, a password wall, or a store that isn't running Shopify.

### Monitor changes on a schedule + webhook

Use monitor mode when you care about **what changed**, not another full catalogue dump. State is stored in a named Apify key-value store across runs. The first run for each store/collection/output-mode watch is a baseline: it emits no billable changes and one free `store_summary` row. Later runs emit only:

- `product_added`
- `product_removed` — only after this run reached the real end of pagination; never inferred from a partial/capped fetch
- `price_changed` — includes previous/current price and percentage change
- `sale_started` / `sale_ended` — based on Shopify compare-at pricing
- `back_in_stock` / `out_of_stock`

Example scheduled input:

```json
{
  "storeUrls": ["https://www.allbirds.com"],
  "outputMode": "variant",
  "maxProductsPerStore": 2000,
  "maxItems": 200,
  "onlyChangesSinceLastRun": true,
  "stateStoreName": "allbirds-competitor-watch"
}
```

A concrete recipe:

1. Run the input above once to establish the baseline.
2. In **Apify Console → Schedules**, create a schedule for this Actor every 6 hours and reuse the same input and `stateStoreName`.
3. In **Integrations → Webhooks**, add an **Actor run succeeded** webhook that POSTs to your endpoint (or an n8n/Make/Zapier webhook URL).
4. In the webhook handler, read the completed run's default dataset. Ignore `store_summary` if you only want alerts; every other row is a billable change event ready to route to Slack, email, a database or a pricing workflow.

Example change row:

```json
{
  "eventType": "price_changed",
  "storeDomain": "www.allbirds.com",
  "productId": "7340901859408",
  "variantId": "42146889039952",
  "title": "Women's Allbirds Flip Flop - Dusty Pink",
  "sku": "A12513W050",
  "previousPrice": 50,
  "currentPrice": 25,
  "priceChange": -25,
  "priceChangePct": -50,
  "observedAt": "2026-09-28T08:00:00.000Z"
}
```

### Use cases

- **Competitor price monitoring**: run it on a schedule against competitor stores and compare `priceMin`/`priceMax`/`onSale` over time.
- **Dropshippers**: pull a supplier or competitor's full catalogue with prices and stock in one export.
- **Market research**: compare vendors, product types, tag usage and pricing across many Shopify stores at once.

### Input

```json
{
  "storeUrls": ["https://www.allbirds.com"],
  "collectionHandle": "",
  "outputMode": "product",
  "onlyInStock": false,
  "maxProductsPerStore": 250,
  "maxItems": 100
}
```

| Field | Type | Default | Description |
|---|---|---:|---|
| `storeUrls` | string\[] | required | 1–100 Shopify store domains or URLs. A bare domain, `http(s)://`, `www.` and any trailing path are all accepted; only the domain is used. |
| `collectionHandle` | string | *(none)* | Limit every store to one collection (e.g. `mens-shoes`, or a full collection URL). Leave empty for the whole catalogue. |
| `outputMode` | `"product"` | `"variant"` | `"product"` | One row per product (with a nested `variants[]`) or one row per variant. |
| `onlyInStock` | boolean | `false` | Deliver only rows with at least one available variant (product mode) or an available variant (variant mode). Filtered-out rows are never charged. |
| `maxProductsPerStore` | integer | 250 | 1–20,000. Cap on distinct products read per store. |
| `maxItems` | integer | 100 | 1–20,000. Normal mode: product/variant rows. Monitor mode: billable change-event rows. Free summary rows do not count. |
| `onlyChangesSinceLastRun` | boolean | `false` | Opt-in monitor mode. First run saves a baseline; later runs output only change events. Off by default, so normal runs are unchanged. |
| `stateStoreName` | string | `shopify-store-products-scraper-state` | Named cross-run state store used only in monitor mode. Reuse the same name for the same watch; change it to start a fresh baseline. |

### Output fields

| Category | Fields |
|---|---|
| Store | `store`, `storeDomain`, `collectionHandle` |
| Identity | `productId`, `handle`, `title`, `url` |
| Catalogue | `vendor`, `productType`, `tags` |
| Timestamps | `createdAt`, `updatedAt`, `publishedAt`, `scrapedAt` |
| Price | `priceMin`, `priceMax`, `compareAtPriceMax`, `onSale`, `currency` |
| Stock | `available`, `variantCount`, `variants[]` (product mode) |
| Variant (variant mode) | `variantId`, `variantTitle`, `sku`, `price`, `compareAtPrice`, `available`, `option1`, `option2`, `option3` |
| Media & copy | `imageUrls` (first 5), `descriptionText` (HTML stripped, truncated to 2,000 characters) |

Missing values are returned as `null`. Nothing is invented.

### Output example (product mode)

```json
{
  "store": "https://www.allbirds.com",
  "storeDomain": "www.allbirds.com",
  "productId": "7340901859408",
  "handle": "womens-allbirds-flip-flop-dusty-pink",
  "title": "Women's Allbirds Flip Flop - Dusty Pink",
  "vendor": "Allbirds",
  "productType": "Shoes",
  "tags": ["shoprunner", "..."],
  "url": "https://www.allbirds.com/products/womens-allbirds-flip-flop-dusty-pink",
  "priceMin": 25,
  "priceMax": 25,
  "compareAtPriceMax": 50,
  "onSale": true,
  "currency": null,
  "available": true,
  "variantCount": 7,
  "variants": [
    { "id": "42146889039952", "title": "5", "sku": "A12513W050", "price": 25, "compareAtPrice": 50, "available": false, "option1": "5", "option2": null, "option3": null }
  ],
  "imageUrls": ["https://cdn.shopify.com/s/files/1/1104/4168/files/A12513_..._LEFT.png"],
  "descriptionText": "Sun on your feet. Comfort underneath. Light, easy, and made for warm weather...",
  "collectionHandle": null,
  "scrapedAt": "2026-09-27T00:00:00.000Z"
}
```

In `outputMode: "variant"`, the same fields above (minus the nested `variants[]`) are repeated on one row per variant, with `variantId`/`variantTitle`/`sku`/`price`/`compareAtPrice`/`option1-3` added and `available` meaning *this variant's* stock status.

### Use it as an API

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("lukehunter/shopify-store-products-scraper").call(
    run_input={
        "storeUrls": ["https://www.allbirds.com", "https://www.gymshark.com"],
        "maxItems": 500,
    }
)

for product in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(product["title"], product["priceMin"], product["available"])
```

### Pricing and cost control

Pay per delivered row (product, or variant in variant mode).

**$0.002 per product — 1,000 products = $2.**

| Rows delivered | Cost |
|---:|---:|
| 20 | $0.04 |
| 1,000 | $2.00 |
| 10,000 | $20.00 |

There is no charge for starting a run, and no charge for a store that is skipped, blocked, or produces nothing. In monitor mode, only change-event rows use the existing `product` charge; each per-store `store_summary` row is pushed free. `maxItems` gives you a hard upper bound on billable rows.

### Limitations

- **Only public product data the store exposes on `products.json`.** No customer data, order data, checkout data, or anything requiring login — this Actor never collects that, by design.
- Some stores **disable** their public `products.json` feed, put it behind a store password, or block automated requests entirely; when that happens the run reports it honestly per store instead of failing the whole run or crashing.
- **`currency` is always `null`.** Shopify's public `products.json` endpoint does not publish a currency anywhere in its response, and this Actor never guesses one from a store's country or price formatting.
- Prices reflect the storefront's default price list at scrape time; **quantity breaks, B2B price lists and localized/multi-currency pricing are not reflected.**
- **Draft, unpublished and password-protected products** never appear in the public feed and so cannot be collected.
- This Actor honours every store's `robots.txt` and never bypasses a password wall, CAPTCHA or block — a blocked store is reported, not evaded.
- This is an independent tool, not affiliated with, endorsed by, or connected to Shopify Inc. or any store scraped. You are responsible for using the collected data in accordance with the relevant store's and Shopify's terms and applicable law.

### FAQ

#### Does this work on any Shopify store?

Any store running the standard Shopify storefront with its public `products.json` feed enabled and not password-protected. A minority of stores disable this feed or block automated requests; those are reported per store, not treated as a run failure.

#### Can I scrape just one collection?

Yes — set `collectionHandle` to the collection's handle (e.g. `mens-shoes`) or paste its full URL.

#### Does it collect customer or order data?

No, never. Only the public product catalogue.

#### Why is `currency` always null?

Shopify's public `products.json` response never states a currency. Guessing one from a country-code TLD or price format would sometimes be wrong, so this Actor reports `null` rather than a guess you can't verify.

### Related Actors

Other data tools from the same developer, built to the same standard: official or public sources, hard cost caps, and honest documentation of limits.

- **[Walmart Category Scraper](https://apify.com/lukehunter/walmart-category-scraper)**: product names, prices, was-prices and ratings from Walmart category pages.
- **[Vinted Scraper](https://apify.com/lukehunter/vinted-scraper)**: Vinted search results with prices, brands, sizes and favourites, across any Vinted country.
- **[AliExpress Search Scraper](https://apify.com/lukehunter/aliexpress-scraper)**: AliExpress search results with prices, discounts, ratings and sold counts, by keyword.
- **[Google Play App Monitor & Scraper](https://apify.com/lukehunter/google-play-scraper)**: ratings, installs, developer contact info and pricing for any Google Play app.
- **[Apple App Store Reviews Scraper](https://apify.com/lukehunter/app-store-reviews-scraper)**: Apple App Store reviews for any iOS app, across countries, with rating, version and date.
- **[Spotify Scraper](https://apify.com/lukehunter/spotify-scraper)**: play counts, monthly listeners and playlist track lists for any public Spotify artist, playlist, album or track.
- **[Hospital Price Transparency Enforcement Leads](https://apify.com/lukehunter/hospital-price-transparency-enforcement-leads)**: hospitals with recent CMS price transparency warning notices, CAP requests and CMP notices.
- **[Hospital Ownership Change Radar](https://apify.com/lukehunter/hospital-chow-radar)**: hospitals that just changed owner, with buyer, seller and effective date from CMS filings.

# Actor input Schema

## `storeUrls` (type: `array`):

Shopify store domains or URLs, e.g. "allbirds.com" or "https://www.allbirds.com". Any path is ignored — this Actor always reads the store's own /products.json. Maximum 100 stores per run.

## `collectionHandle` (type: `string`):

Limit every store to one collection instead of the whole catalogue, e.g. "mens-shoes" for https://store.com/collections/mens-shoes (a full collection URL also works). Leave empty to scrape each store's entire product catalogue.

## `outputMode` (type: `string`):

"product": one row per product, with a nested variants\[] array. "variant": one row per variant, with the parent product's fields repeated on each row.

## `onlyInStock` (type: `boolean`):

Normal scrape mode only: deliver only in-stock products/variants. Monitor mode always observes both in-stock and out-of-stock rows so back\_in\_stock/out\_of\_stock changes can be detected.

## `maxProductsPerStore` (type: `integer`):

Cap on distinct products read from any single store, 1-20000, default 250. In variant mode this still limits distinct products, not variant rows.

## `maxItems` (type: `integer`):

Hard cap on delivered rows, 1-20000, default 100. In normal mode these are product/variant rows; in monitor mode these are billable change-event rows (free summary rows do not count).

## `onlyChangesSinceLastRun` (type: `boolean`):

Off by default. Turn on for scheduled monitoring: the first run saves a baseline in a named key-value store and outputs only a free store\_summary row; later runs output only product/variant change events and charge only those change rows. Default one-off scraping is unchanged.

## `stateStoreName` (type: `string`):

Only used when change monitoring is on. Named Apify key-value store that persists each store/collection/output-mode snapshot across scheduled runs. Keep the same name for the same watch; use a different name to start an independent baseline. Letters, numbers and hyphens only.

## Actor input object example

```json
{
  "storeUrls": [
    "https://www.allbirds.com"
  ],
  "outputMode": "product",
  "onlyInStock": false,
  "maxProductsPerStore": 250,
  "maxItems": 20,
  "onlyChangesSinceLastRun": false,
  "stateStoreName": "shopify-store-products-scraper-state"
}
```

# Actor output Schema

## `products` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "storeUrls": [
        "https://www.allbirds.com"
    ],
    "collectionHandle": "",
    "outputMode": "product",
    "onlyInStock": false,
    "maxProductsPerStore": 250,
    "maxItems": 20,
    "onlyChangesSinceLastRun": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("lukehunter/shopify-store-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "storeUrls": ["https://www.allbirds.com"],
    "collectionHandle": "",
    "outputMode": "product",
    "onlyInStock": False,
    "maxProductsPerStore": 250,
    "maxItems": 20,
    "onlyChangesSinceLastRun": False,
}

# Run the Actor and wait for it to finish
run = client.actor("lukehunter/shopify-store-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "storeUrls": [
    "https://www.allbirds.com"
  ],
  "collectionHandle": "",
  "outputMode": "product",
  "onlyInStock": false,
  "maxProductsPerStore": 250,
  "maxItems": 20,
  "onlyChangesSinceLastRun": false
}' |
apify call lukehunter/shopify-store-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lukehunter/shopify-store-products-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UZMcrT3DC94t7mqLm/builds/AmH100xKK3m9RC182/openapi.json
