# Product Page Scraper: Price, Stock, Brand from Any Shop (`swiftkit/product-pages`) Actor

Extract product name, price, currency, stock status, brand, SKU, GTIN, rating and image from product pages of almost any online shop, using the structured data shops publish for Google (schema.org). Category pages, price-change tracking for monitoring. $2 per 1,000 products.

- **URL**: https://apify.com/swiftkit/product-pages.md
- **Developed by:** [SwiftKit](https://apify.com/swiftkit) (community)
- **Categories:** E-commerce, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 products

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Product Page Scraper: Price, Stock, Brand from Any Shop

Give it product pages from **almost any online shop** and get clean product rows: **name, price,
currency, stock status, brand, SKU, GTIN/barcode, MPN, rating, review count, image, category,
description and seller**. Turn on **price tracking** and schedule it to see what got cheaper, more
expensive or went out of stock since the last run.

How it works: most shops publish **structured product data** in their pages (schema.org JSON-LD,
the same data Google uses for shopping results). The tool reads that first, then microdata, then
Open Graph product tags. No site-specific scraping, so it works across Shopify, WooCommerce,
Magento, BigCommerce, custom shops and big retailers that publish the data.

- **Category pages:** turn on "Find product links" and paste a category or search page; product
  links on it are collected and extracted
- **Price tracking:** previous price, change, % change and "availability changed" per product
- **Polite by design:** respects robots.txt, a few pages at a time, and stops at bot checks or
  waiting-room pages instead of trying to get around them

### Who it's for

- **Ecommerce teams:** watch competitors' prices and stock on the products that matter.
- **Brands and distributors:** check how retailers price and list your products (GTIN matching).
- **Affiliates and comparison sites:** keep prices and availability up to date.
- **Analysts:** price and rating snapshots across shops.

### Input

| Option | Default | What it does |
|---|---|---|
| Product or category pages | – | URLs from any shop |
| Find product links on category pages | off | Follow product links on listing pages |
| Max products | 200 | Stop after this many |
| Track price changes | off | Compare with the previous run |
| Price history name | `default` | Separate histories |
| Include description | on | Plain-text description |
| Respect robots.txt | on | Skip disallowed pages |
| Pages at once | 4 | Concurrency |

### Output

A real result:

```json
{
  "url": "https://www.zappos.com/p/crocs-classic-clog-black/product/7153812/color/3",
  "status": "ok",
  "name": "Classic Clogs",
  "brand": "Crocs",
  "price": 49.99,
  "currency": "USD",
  "availability": "InStock",
  "rating": 5,
  "source": "json-ld",
  "site": "zappos.com"
}
```

With price tracking on, rows also have `previousPrice`, `previousCheckAt`, `priceChange`,
`priceChangePct` and `availabilityChanged`.

| status | Meaning |
|---|---|
| `ok` | Product extracted. |
| `no_product_data` | The page has no structured product data. Not charged. |
| `blocked_by_bot_protection` | The shop showed a bot check or waiting room. Not charged. |
| `blocked_by_robots_txt` | The shop asks bots not to visit. Not charged. |
| `unreachable` | The page didn't load (e.g. 403/404). Not charged. |

### Use it from your AI agent (MCP)

Claude, Cursor and other MCP clients can call this tool directly through Apify's MCP server. Add it as
an MCP server / connector:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=swiftkit/product-pages",
      "headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
    }
  }
}
```

Then just ask, for example: *"What do these three product pages charge, and are they in stock? …"*. Each call is billed like a normal run.

### Pricing

You pay **per product extracted**. Every other status is free. See the Pricing tab.

### Limits, honestly

- Works when the shop publishes structured data in its HTML (most do, because Google rewards it).
  Shops that build pages only with JavaScript, or that block bots (many large retailers), return
  `no_product_data` or `blocked`. That's the shop's choice, and the tool respects it.
- When a page lists several variants, the price is the **lowest** variant price (and `priceMax` the highest).
- Shops may show prices in the visitor's local currency. Runs come from Apify's servers (mostly the
  US), so you usually see USD for international shops.
- Reviews' text and reviewer names are not collected.
- You're responsible for how you use the data; respect each shop's terms.

### More tools from SwiftKit

- [Shopify Store Products](https://apify.com/swiftkit/shopify-products): every product of a Shopify store in one run
- [Website Change Monitor](https://apify.com/swiftkit/website-monitor): watch any page for text changes
- [Tech Stack Detector](https://apify.com/swiftkit/tech-stack): which ecommerce platform a shop runs on

### Questions?

Open an issue on the Issues tab.

# Actor input Schema

## `urls` (type: `array`):

Product page URLs from any shop. With "Find product links" on, category or search pages work too.

## `findProductLinks` (type: `boolean`):

When a page isn't a product page, collect links on it that look like product pages (same site) and extract those.

## `maxProducts` (type: `integer`):

Stop after this many products.

## `trackPrices` (type: `boolean`):

Remember each product's price between runs and add previous price, change and % change, plus whether availability changed. Schedule it for price monitoring.

## `monitorName` (type: `string`):

Keeps separate price histories (e.g. competitors, suppliers). Use the same name every run.

## `includeDescription` (type: `boolean`):

Product description as plain text (up to 5,000 characters).

## `respectRobotsTxt` (type: `boolean`):

Skip pages the shop asks bots not to visit.

## `maxConcurrency` (type: `integer`):

Keep it low to be gentle with shops.

## Actor input object example

```json
{
  "urls": [
    "https://www.allbirds.com/products/mens-tree-runners",
    "https://www.zappos.com/p/crocs-classic-clog-black/product/7153812/color/3"
  ],
  "findProductLinks": false,
  "maxProducts": 200,
  "trackPrices": false,
  "monitorName": "default",
  "includeDescription": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

One row per product page.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.allbirds.com/products/mens-tree-runners",
        "https://www.zappos.com/p/crocs-classic-clog-black/product/7153812/color/3"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("swiftkit/product-pages").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.allbirds.com/products/mens-tree-runners",
        "https://www.zappos.com/p/crocs-classic-clog-black/product/7153812/color/3",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("swiftkit/product-pages").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.allbirds.com/products/mens-tree-runners",
    "https://www.zappos.com/p/crocs-classic-clog-black/product/7153812/color/3"
  ]
}' |
apify call swiftkit/product-pages --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,swiftkit/product-pages"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HPLg3KAQkJiC6jEYc/builds/3x073akKllLdhhZq6/openapi.json
