# WooCommerce Products Scraper - Prices, Variations & Stock (`tinyrex/woocommerce-products-scraper`) Actor

Export all products from any WooCommerce store: prices, sale prices, variations with SKU and stock, categories, brands, images. Monitors new products and price changes. Pay per product.

- **URL**: https://apify.com/tinyrex/woocommerce-products-scraper.md
- **Developed by:** [TinyRex](https://apify.com/tinyrex) (community)
- **Categories:** E-commerce, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 products

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## WooCommerce Products Scraper: prices, variations, stock and monitoring

Export every product from any WooCommerce store: prices, sale prices, variations (size, color) with their own SKU, price and stock, categories, tags, brands, ratings, images and descriptions. Add store domains, category URLs or product URLs and get clean, structured data. You pay only for the products you get.

**Why this one:** **$0.60 per 1,000 products**, variation details included, while popular WooCommerce scrapers on the Store charge about $0.90–3 per 1,000. Non-WooCommerce or blocked sites and filtered-out products are **free**.

### Why this scraper

- **Fast and light.** It reads the public WooCommerce **Store API** that powers the store's own cart and product blocks. No browser, so runs are quick and cheap.
- **Full variation data.** Price, regular price, SKU and stock text for every variation, nested in the product or as flat rows. No extra cost.
- **Clean prices.** Decimal prices in the store currency (with currency code and symbol), sale price and discount %.
- **Monitoring built in.** Get only **new products**, or new products plus **price and stock changes** with the previous values, on every scheduled run.
- **Pay only for results.** Sites that are not WooCommerce, blocked sites and products removed by your filters are **free**.
- **Respectful by design.** The scraper follows each store's `robots.txt`. If a store asks bots not to use its API, it is skipped (and free).

### What you get

| Field | Example |
|---|---|
| Store name, URL | Nalgene, https://nalgene.com |
| Product ID, SKU, slug, URL, title, type | 905247, 342021, 32oz-wide-mouth-sustain, …/product/32oz-wide-mouth-sustain/, 32oz Wide Mouth Sustain, variable |
| Price, max price, regular price, sale price, discount % | 17.99, 21.99, 24.99, 19.99, 20 |
| Currency | USD, $ |
| Stock | in stock, on backorder, stock status, stock text ("12 in stock"), low stock remaining |
| Reviews | average rating, review count |
| Categories, tags, brands | \[Water Bottles], \[bpa-free], \[Nalgene] |
| Attributes | Color: Blue, Gray … (marks which ones are variation options) |
| Variations | id, title, attributes, SKU, price, regular price, sale price, on sale, in stock, stock text, image, URL |
| Images | main image + all image URLs |
| Weight, dimensions | as set in the store |
| Description | short and full description, plain text or HTML |
| In monitoring mode | change type, previous price, previous stock status |

### How to use

1. Add stores to **WooCommerce stores, categories or products**:
   - a domain: `nalgene.com`
   - a category: `https://nalgene.com/product-category/accessories/` (custom category URLs work too)
   - a single product: `https://nalgene.com/product/four-wheeling-sticker/`
2. Optional: a **search query** (uses the store's own search, fastest for big stores), keywords, exclude keywords, categories, brands, price range, only in stock, only on sale.
3. Choose **one row per product** (variations nested) or **one row per variation** (flat, best for spreadsheets and price tracking).
4. Run, then download the results as JSON, CSV, Excel or HTML, or use them through the API, webhooks or integrations (Google Sheets, Make, Zapier, n8n, Slack).

#### Example input

```json
{
  "stores": ["https://nalgene.com", "https://www.mandarinstone.com/collections/all-terrazzo/"],
  "outputMode": "products",
  "onlyInStock": true,
  "maxProductsPerStore": 500
}
```

#### Example output (shortened)

```json
{
  "storeName": "Mandarin Stone",
  "productId": 499641,
  "url": "https://www.mandarinstone.com/product/terrazzo-ivory-honed/",
  "title": "Terrazzo Ivory Honed",
  "type": "variable",
  "price": 0.66,
  "maxPrice": 13.68,
  "currency": "GBP",
  "inStock": true,
  "categories": ["All Terrazzo", "New arrivals"],
  "attributes": [{ "name": "Size", "values": ["200x200x12", "200x50x12", "406x406x18"], "usedForVariations": true }],
  "variationsCount": 4,
  "variations": [{ "id": 500195, "title": "Size: 200x50x12", "sku": "TCTEIVOSC20X5", "price": 0.66, "inStock": true, "stockText": "4094 in stock" }],
  "featuredImage": "https://www.mandarinstone.com/app/uploads/2026/08/Terrazzo-Ivory-Honed-scaled.jpg"
}
```

### Monitoring new products and price changes

Set **Monitoring mode** and schedule the Actor (for example daily):

- **Only new products:** each run returns products that were not seen before.
- **New products and price/stock changes:** also returns products whose price or stock status changed, with `changeType`, `previousPrice` and `previousInStock`.

The first monitoring run returns all matching products as the baseline. Use a different **State key** for each schedule or filter set. The state is stored in your account in the key-value store `woocommerce-products-state`.

### Run report

Besides the dataset, each run saves:

- **`STORES`**: one record per input with the status (`ok`, `not_woocommerce`, `blocked`, `disallowed_by_robots`, `unreachable`, `not_found`, `error`), store name, currency, API used and product counts.
- **`SUMMARY`**: totals for the run.

### Pricing

Pay per event:

- **Product** (`product`): one charge per row saved to your dataset (per product, or per variation in variation mode). Variation details in product mode are included.
- **Free:** sites that are not WooCommerce, blocked sites, stores whose robots.txt disallows the API, unreachable sites and filtered-out products.

See the *Pricing* tab for current prices. Set a maximum cost per run and the Actor stops gracefully when it is reached.

### Tips

- Use **variation mode** with **only on sale** to build a discount feed.
- Paste a **category URL** to scrape just one category of a big store, or use **search query**.
- Find WooCommerce stores first with a technology detector, then pass the domains here.
- If a store is slow or rate-limits (HTTP 429), lower **Stores in parallel** or use Apify Proxy (Advanced section).

### Limitations

- Works with WooCommerce stores whose public Store API is available (WooCommerce 6+; most stores). Some stores disable it, protect it with bot checks, or disallow it in robots.txt; these are skipped and not charged.
- Only data the store shows publicly is available. Sales numbers, cost prices and exact stock quantities (unless the store shows them) are not.
- The Store API does not expose product dates; use monitoring mode to detect new products.
- Prices are shown as the store shows them to a guest visitor (tax display and currency follow the store settings).

### Use with AI agents (MCP)

This Actor works as a tool for AI agents through the **Apify MCP server** at `https://mcp.apify.com`. Connect Claude, Cursor, VS Code, n8n or any other MCP client to `https://mcp.apify.com?tools=tinyrex/woocommerce-products-scraper` and the agent can run it from a plain-language request, for example: *"Get all products and variation prices from nalgene.com"*. Results come back as clean, structured JSON at the same pay-per-result price, and failed inputs stay free.

📘 **Step-by-step guide with Python, JavaScript, curl and MCP examples:** [How to export products from any WooCommerce store to CSV or JSON](https://emirmrkaljevic.github.io/tinyrex-data-tools/export-woocommerce-store-products-to-csv/)

### Related actors

- [Tech Stack Detector](https://apify.com/tinyrex/tech-stack-detector): Have a list of domains? Find out which ones run WooCommerce first (use its `onlyIfUses` lead filter with `WooCommerce`), then feed them here.
- [Shopify Products Scraper](https://apify.com/tinyrex/shopify-products-scraper): The same product export for Shopify stores, including headless ones.
- [ATS Jobs Scraper](https://apify.com/tinyrex/ats-jobs-scraper): See which companies are hiring: jobs from Greenhouse, Lever, Ashby, Personio and more.
- [Workday Jobs Scraper](https://apify.com/tinyrex/workday-jobs-scraper): Jobs from enterprise Workday career sites.

### FAQ

**Is it legal?** The Actor reads public product data through the Store API that WooCommerce provides for storefronts. It does not log in, does not access customer or order data, and respects robots.txt. Use the data in line with the store's terms and your local law.

**How fast is it?** In our test, 24 stores with 2,637 products took about 4 minutes; most of the time is spent waiting for the stores' servers. Speed depends on store size and hosting.

**My store isn't recognized.** Check the `STORES` record for the reason. If the status is `not_woocommerce` but the store runs WooCommerce, open an issue with the URL and we will take a look.

# Actor input Schema

## `stores` (type: `array`):

Store domains (nalgene.com), store URLs, category URLs (https://store.com/product-category/mugs/) or product URLs (https://store.com/product/blue-mug/). Uses the store's public WooCommerce Store API. Sites that are not WooCommerce, block access or disallow the API in robots.txt are skipped and not charged.

## `maxProductsPerStore` (type: `integer`):

Maximum number of products to save per input (after filters). 0 = all products. Newest products come first.

## `outputMode` (type: `string`):

Product: one row per product, variations nested. Variation: one flat row per variation (size/color) - easiest for spreadsheets and price tracking; simple products still give one row. One charge per row.

## `search` (type: `string`):

Optional. Uses the store's own product search, so only matching products are downloaded - fastest way to scan big stores for one product line.

## `keywords` (type: `array`):

Keep only products whose title, SKU, categories, tags or brands contain any of these words (case-insensitive).

## `excludeKeywords` (type: `array`):

Drop products whose title or tags contain any of these words.

## `categories` (type: `array`):

Keep only products in these categories (name or slug, partial match). To download just one category, you can also paste its URL as input.

## `brands` (type: `array`):

Keep only products of these brands (partial match). Works on stores that use WooCommerce Brands.

## `minPrice` (type: `integer`):

Lowest price must be at least this (store currency). 0 = no limit.

## `maxPrice` (type: `integer`):

Lowest price must be at most this (store currency). 0 = no limit.

## `onlyInStock` (type: `boolean`):

Keep only products that are in stock.

## `onlyOnSale` (type: `boolean`):

Keep only discounted products.

## `includeVariations` (type: `boolean`):

Add price, regular price, SKU and stock for every variation of variable products. Costs nothing extra (one request per 100 variations).

## `descriptionFormat` (type: `string`):

Short and full description as plain text, HTML, both, or none (smaller output).

## `monitorMode` (type: `string`):

Off: return all matching products. New: only products not seen in previous runs. Changes: new products plus price or stock changes (with previous values). State is kept per store under the state key. The first monitoring run returns everything as the baseline.

## `stateKey` (type: `string`):

Name of the monitoring state. Use different keys for different schedules or filters.

## `maxConcurrency` (type: `integer`):

How many stores are processed at the same time. Pages of one store are fetched politely one after another.

## `proxyConfiguration` (type: `object`):

Usually not needed. Use a proxy if a store rate-limits requests (HTTP 429).

## Actor input object example

```json
{
  "stores": [
    "https://nalgene.com",
    "mandarinstone.com"
  ],
  "maxProductsPerStore": 100,
  "outputMode": "products",
  "search": "",
  "keywords": [],
  "excludeKeywords": [],
  "categories": [],
  "brands": [],
  "minPrice": 0,
  "maxPrice": 0,
  "onlyInStock": false,
  "onlyOnSale": false,
  "includeVariations": true,
  "descriptionFormat": "text",
  "monitorMode": "off",
  "stateKey": "default",
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `products` (type: `string`):

Products with prices, variations, stock, categories and images.

## `stores` (type: `string`):

For every input: status (ok, not_woocommerce, blocked, disallowed_by_robots, not_found...), store name, currency, products found/saved, notes.

## `summary` (type: `string`):

Totals for the run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "stores": [
        "https://nalgene.com",
        "mandarinstone.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinyrex/woocommerce-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "stores": [
        "https://nalgene.com",
        "mandarinstone.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tinyrex/woocommerce-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "stores": [
    "https://nalgene.com",
    "mandarinstone.com"
  ]
}' |
apify call tinyrex/woocommerce-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinyrex/woocommerce-products-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DOrw9o0FzP8gPpxdQ/builds/2dq71BofXrHrAGbsj/openapi.json
