# Ecommerce Product Page Scraper: Price, Stock, GTIN from URLs (`pistachio_implementation/product-page-extractor`) Actor

Paste product page URLs from almost any online store and get name, price, currency, sale price, stock status, brand, SKU, GTIN, rating, images and variants, read from the Schema.org and Open Graph data stores publish. Plain HTTP, no browser. Blocked pages are free.

- **URL**: https://apify.com/pistachio\_implementation/product-page-extractor.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** E-commerce, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.50 / 1,000 product extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Ecommerce Product Page Scraper: Price, Stock, GTIN from URLs

Paste links to product pages from almost any online store and get one clean row per product: name, brand, price, currency, list price and sale flag, stock status, SKU, GTIN (EAN or UPC), MPN, rating, review count, images and every variant the page lists.

It works on Shopify, WooCommerce, BigCommerce, Magento, Wix, Squarespace, PrestaShop and custom stores, and on large retailers that let ordinary visitors in, because it reads the **Schema.org product data** (`JSON-LD` and microdata) and **Open Graph product tags** that stores publish for Google Shopping and social previews. No per site setup, no browser, no AI guessing.

### What you can use it for

- **Price monitoring:** schedule a list of competitor product pages daily and track price, sale price and stock.
- **Catalog enrichment:** fill in GTINs, brands, images and descriptions for products you resell.
- **MAP and reseller checks:** see what each reseller charges for the same GTIN.
- **Market research:** compare prices, ratings and review counts across stores.
- **AI agents:** one job, one input (a list of URLs), a predictable schema.

### How it works

1. Fetches each page with a plain HTTP request (one request at a time per site, with a pause between requests).
2. Reads `JSON-LD` `Product` and `ProductGroup` blocks first (including `@graph`, `AggregateOffer`, list price specifications and `hasVariant`), then fills gaps from microdata, then from Open Graph `product:` tags.
3. Pages that are blocked, missing or have no product data are written with a clear `status` and **are not charged**.

It respects robots.txt by default (including rules that name Apify), never logs in, never uses cookies from an account, never solves captchas and never uses proxies to get around blocks.

### Input

| Field | What it does | Default |
|---|---|---|
| Product page URLs | Links to single product pages, one per line | required |
| Include description | Plain text description, up to 5,000 characters | on |
| Include variants | Size, color, SKU, GTIN, price and stock per variant | on |
| Respect robots.txt | Skip pages the site closes to automated tools | on |
| Parallel requests | Pages fetched at once across different sites | 5 |
| Pause between requests to the same site | Politeness delay in milliseconds | 1000 |
| Maximum URLs | Stop after this many URLs | 1000 |

Example input:

```json
{
  "productUrls": [
    "https://www.allbirds.com/products/mens-tree-runners",
    "https://www.ikea.com/us/en/p/billy-bookcase-white-00263850/"
  ],
  "includeDescription": true,
  "includeVariants": true
}
```

### Output

One row per URL. Example (trimmed):

```json
{
  "url": "https://www.allbirds.com/products/mens-tree-runners",
  "status": "ok",
  "dataSources": ["json-ld", "open-graph"],
  "domain": "allbirds.com",
  "name": "Men's Tree Runner",
  "brand": "Allbirds",
  "price": 100,
  "currency": "USD",
  "listPrice": null,
  "onSale": null,
  "lowPrice": 100,
  "highPrice": 100,
  "availability": "InStock",
  "inStock": true,
  "sku": "MENS_TREE_RUNNERS",
  "gtin": null,
  "rating": null,
  "reviewCount": null,
  "imageUrl": "https://cdn.shopify.com/s/files/1/1104/4168/files/...png",
  "variantCount": 49,
  "variants": [
    { "name": "Men's Tree Runner, Jet Black (White Sole), 8", "sku": "...", "gtin": "...", "size": "8", "price": 100, "currency": "USD", "availability": "InStock", "inStock": true }
  ],
  "scrapedAt": "2026-09-27T12:00:00.000Z"
}
```

`status` is one of:

| Status | Meaning | Charged |
|---|---|---|
| `ok` | Product data found | yes |
| `no_product_data` | Page loaded but has no usable product data (no price and no Schema.org product) | no |
| `blocked` | The site answered with a bot check or access denied page | no |
| `http_error` | 404, 410, 500 and similar | no |
| `skipped_robots_txt` | robots.txt closes the page to automated tools | no |
| `error` | Timeout or network failure | no |

A `RUN_SUMMARY` record in the run's key value store counts each status.

### Pricing

Pay per event. No subscription and no platform usage charge on top.

| Event | Price |
|---|---|
| Product extracted | $0.0015 (1.50 dollars per 1,000 products) |

Blocked pages, pages without product data, errors and robots.txt skips are free. Example: tracking 500 competitor products every day costs about $0.75 a day. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- **Sites with strong bot protection** (for example Amazon, Walmart, Best Buy, Target, Home Depot and others using Akamai, PerimeterX, DataDome or Cloudflare challenges) often answer datacenter traffic with a block page. Those rows come back as `blocked` and cost nothing. This actor does not try to get around blocks.
- Only what the page publishes as structured data is returned. If a store leaves out GTIN or stock in its Schema.org data, those fields are empty.
- Prices are the ones shown to a visitor from a US datacenter with no cookies. Stores that change price or currency by country may show different values to you.
- Prices rendered only by JavaScript after the page loads, with no structured data, are not seen (no browser is used).
- One product per URL. Category and search pages are not crawled; use a sitemap or catalog tool to collect product URLs first.

### FAQ

**Which stores work best?** Shopify, WooCommerce, BigCommerce, Wix, Squarespace and most modern stores publish complete Schema.org product data because Google Shopping needs it.

**Do I pay for pages that fail?** No. Only rows with status `ok` are charged.

**Can I use it for price alerts?** Yes. Schedule it, then compare `price`, `listPrice` and `inStock` between runs in your sheet, database or automation tool.

**Does it use AI?** No. It reads the data the store itself publishes, so results are exact and repeatable.

**Is it affiliated with any store?** No. It is an independent tool that reads public product pages.

# Actor input Schema

## `productUrls` (type: `array`):

Links to individual product pages, one per line, from any store (Shopify, WooCommerce, BigCommerce, Magento, Wix, Squarespace, custom sites and large retailers that publish Schema.org product data).

## `includeDescription` (type: `boolean`):

Add the product description as plain text (up to 5,000 characters).

## `includeVariants` (type: `boolean`):

Add the list of variants (size, color, SKU, GTIN, price, stock) when the page publishes them.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt closes to automated tools. Skipped pages are free.

## `maxConcurrency` (type: `integer`):

How many pages to fetch at the same time across different sites. Each single site still gets one request at a time.

## `delayPerSiteMs` (type: `integer`):

Politeness delay for pages on the same website.

## `maxItems` (type: `integer`):

Process at most this many URLs from the list.

## Actor input object example

```json
{
  "productUrls": [
    "https://www.allbirds.com/products/mens-tree-runners",
    "https://www.ikea.com/us/en/p/billy-bookcase-white-00263850/"
  ],
  "includeDescription": true,
  "includeVariants": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 5,
  "delayPerSiteMs": 1000,
  "maxItems": 1000
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "productUrls": [
        "https://www.allbirds.com/products/mens-tree-runners",
        "https://www.ikea.com/us/en/p/billy-bookcase-white-00263850/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/product-page-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "productUrls": [
        "https://www.allbirds.com/products/mens-tree-runners",
        "https://www.ikea.com/us/en/p/billy-bookcase-white-00263850/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/product-page-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "productUrls": [
    "https://www.allbirds.com/products/mens-tree-runners",
    "https://www.ikea.com/us/en/p/billy-bookcase-white-00263850/"
  ]
}' |
apify call pistachio_implementation/product-page-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/product-page-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TOJbbeQClya4sWQRS/builds/dRgEY4hANjqCaOmus/openapi.json
