# Shopify Products Scraper | Variants & Inventory (`peerless_columbine/shopify-products-inventory-scraper`) Actor

Export Shopify products, variants, prices, images and publicly exposed inventory from known stores. Includes coverage summaries. $1 per 1,000 products plus platform usage. Independent third-party tool.

- **URL**: https://apify.com/peerless\_columbine/shopify-products-inventory-scraper.md
- **Developed by:** [tingyou333 zhuang](https://apify.com/peerless_columbine) (community)
- **Categories:** E-commerce
- **Stats:** 2 total users, 1 monthly users, 60.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 products

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Shopify Products Scraper | Variants & Inventory

Export the public product catalogue of Shopify stores you already know. Each dataset row represents one product, with nested variants, decimal prices, named options, images and availability. Optional detail requests add publicly exposed inventory quantities and product metadata.

This is an independent third-party tool. It is not developed, endorsed or supported by [Shopify](https://www.shopify.com), or by any store being read. It does not discover merchants across the web or provide a database of all Shopify stores.

### Quick start

Click **Try for free**, review the input and run a small export:

```json
{
  "domains": [
    "deathwishcoffee.com"
  ],
  "maxProducts": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "pageSize": 3,
  "maxRunSeconds": 90
}
```

The Console trial and automatic test use a small three-product request with Apify Residential Proxy. Proxy traffic is billed separately, in addition to the product fee and platform usage. Set `proxyConfiguration.useApifyProxy` to `false` to try direct access; omitting the proxy field in JSON also leaves it disabled. Direct cloud requests to the example store returned HTTP 429 on October 1, 2026. Access restrictions are reported as failures, and no sample data is substituted.

Each dataset row is one product with its nested variants. Download JSON for nested data, or select the Products / Nested variants dataset view before exporting.

For the full available catalog, set `maxProducts` to `0` and raise `maxRunSeconds` for the job. Zero is the runtime default; the Console form prefills 3 for a small trial. Positive limits apply separately to each store. Keep the Actor timeout above `maxRunSeconds`; the published default timeout is 3,600 seconds. A time, source, page or fee limit can still make the result partial.

### Pricing

**$1 per 1,000 stored products ($0.001 each), plus Apify platform usage.** A product and all its nested variants count as one result. There is no Actor start fee, per-variant fee or separate inventory-enrichment fee. Failed/skipped products and summary records incur no result fee. Compute, storage, transfer and any proxy you select are additional platform costs. An optional copy to `datasetId` does not create a second product result fee, but storage usage can apply. Use the run spending limit and input limits to control your job.

### Inputs

| Field | Type and default | Actual behavior |
|---|---|---|
| `domains` | required string array; form example `deathwishcoffee.com` | Up to 100 known public store domains. Accepts bare hosts and HTTP(S) URLs. Paths and query filters are ignored. An explicitly empty array uses deathwishcoffee.com. Duplicate normalized www/non-www inputs are skipped. |
| `maxProducts` | integer, `0` | Stored-product limit per store; 0 continues to source exhaustion, subject to disclosed safety/time/budget stops. |
| `includeInventoryDetails` | boolean, `false` | Reads matching product `.json`, Ajax `.js`, and HTML; merges actual public detail fields only. |
| `proxyConfiguration` | object, `{"useApifyProxy":false}` | Apify Proxy groups/country or user-supplied `proxyUrls` participate in request routing. Custom URLs must use HTTP CONNECT. Direct mode remains the runtime default when this field is omitted; the form prefills Residential Proxy with separate costs. |
| `datasetId` | optional 17-character alphanumeric string | Select an existing dataset in the resource picker or supply its ID in JSON. Verifies access, then copies each default-dataset product to it. Empty or omitted uses only the default dataset; the same ID as the default dataset is not written twice. Requires SDK execution, not the standalone local JSONL mode. |
| `runId` | optional 17-character alphanumeric string | Copied as `runId` onto every row. No lookup or upstream run start. Empty means no field. |
| `locale` | string, empty | Explicit locale route, for example `fr` or `en-gb`; the shop must support it. Regional store domains can be supplied directly. No translation or currency conversion. |
| `pageSize` | integer, `250`; form prefill `3`, range 1–250 | Fixed request size on every page, independent of the total product limit. |
| `maxPagesPerStore` | integer, `10000` | Emergency page ceiling. Reaching it means partial collection. |
| `requestTimeoutSecs` | integer, `25` | Timeout per request attempt; allowed 5–120. |
| `maxRetries` | integer, `2` | Bounded retries for transient transport failures, 429 and selected 5xx; maximum 4. No retry on 401/403. |
| `requestIntervalSecs` | number, `0.5` | Minimum delay between sequential source requests; allowed 0.2–10. No concurrent store requests. |
| `maxRunSeconds` | integer, `3300`; form prefill `90` | Graceful collection deadline. Set below the platform run timeout so the summary can be saved. |

Unknown fields and invalid types produce an input error; they are not silently accepted as aliases. Collection, keyword, vendor and price filters are not supported; URL paths do not filter the catalog.

### Output

Standard fields include `store`, `productId`, `handle`, `title`, `vendor`, `url`, `featuredImage`, `imageUrls`, `imageCount`, `imageAltTexts`, `images`, `currency`, `priceMin`, `priceMax`, `compareAtPrice`, `onSale`, `available`, `fullyOutOfStock`, `requiresShipping`, `weightAndUnit`, `variantCount`, `variants`, `options`, `productOptions`, `productType`, `tags`, `description`, `descriptionHtml`, `createdAt`, `publishedAt`, and `updatedAt`.

Prices are decimal monetary values from the catalogue, or from the complete Storefront MoneyV2 set when recovering variants. `priceMin` and `priceMax` span the resulting verified variant set. `compareAtPrice` is the highest exposed compare-at price or null; `onSale` requires a compare-at price higher than that variant's actual price. Ajax integer price units never overwrite catalogue prices. Catalogue currency comes from public metadata, with a public cart fallback; a final response must match the store and locale context. A catalogue redirect to a different origin or locale triggers a fresh currency lookup. If neither confirms a code it is null. Recovery replaces the entire variant price family with MoneyV2 amounts and their consistent currencyCode, recorded as `source.priceContext="storefront_default_market"`; these are the API default market prices, which need not match a requested locale or customer-specific price. No personal cart is used or modified.

Variants preserve numeric IDs, SKU, option-name/value mapping, availability, shipping and taxable flags, prices, weight/grams, image metadata and timestamps. Images preserve IDs, dimensions, positions, linked variant IDs and alt text. `weightAndUnit` includes the first variant's explicitly exposed grams as `{"grams":125,"value":125,"unit":"g"}`. `grams`, `value` and `unit` describe the same exposed weight. During Storefront recovery, grams missing from the catalogue can be calculated from the same variant’s public weight and declared unit; that conversion is recorded in `source.variantRecovery`. If neither source provides a usable weight, the object is null. New variants have null lifecycle timestamps when the public recovery API does not expose them.

Detail mode can add:

- Product: `detailLevel`, `inventoryAccuracy`, `inventoryCount`, `sellingPlanGroups`, `sellingPlanCurrency`, `brand`, `seoTitle`, `seoDescription`, `metafields`, and explicit source values for `status`, `publishedScope` or `templateSuffix`.
- Variant: `inventoryQuantity`, `inventoryManagement`, `inventoryPolicy`, `barcode`, validated `gtin`, `quantityRule`, `requiresSellingPlan`, `sellingPlanAllocations` and `metafields`.

Enrichment-only fields are absent in standard mode. They also remain absent when no detail content was usable. `inventoryAccuracy="exact"` requires one matching public detail response to supply integer quantities for the entire resolved variant set. Quantities from separate partial responses are never combined into an exact total. Variant recovery alone does not recover private quantities. It describes source-reported values at observation time, not independently audited warehouse stock. Quantities may be zero or negative. `inventoryCount` is omitted when any quantity is missing; no partial sum is presented as a total. Barcode and GTIN are distinguished: only a valid numeric GTIN with a matching check digit becomes `gtin`.

Nonempty Ajax selling-plan groups and allocations are retained unchanged only after matching locale `cart.js` confirms the same currency as the product row. `sellingPlanCurrency` and `source.sellingPlanMoney` identify that context and unchanged Shopify Ajax amount format. Unverified or different-currency plans are omitted with a source diagnostic.

An enriched product with other useful details but no complete stock quantities has `inventoryAccuracy="unavailable"`. Brand comes from identified structured Product metadata, never from an assumed vendor mapping. HTML metadata requires an exact matching final product URL and canonical or structured product identity. SEO title and description collapse runs of whitespace to one space and trim surrounding whitespace on both JSON and HTML extraction paths; the wording remains unchanged. Missing metafields, SEO or private publication data are not fabricated.

Added `observedAt` and `source` fields expose provenance, currency origin, exact quantity sources, optional HTML detail URL, requested locale and a possible variant-truncation warning. Use the Products and Nested variants dataset views to inspect these fields.

### Pagination, failure and completeness

The primary route is `/products.json`. A 404, non-JSON response or incompatible shape can lead to `/collections/all/products.json`; this fallback is explicitly labelled because a merchant may curate that collection. Explicit access denials (401/403/430), password gates and identified challenges stop the current store immediately, including during variant recovery or optional detail enrichment; later handles or detail routes are not tried against that block. Existing rows remain in the dataset and the store is partial. Short pages are not assumed to mean the end: the reader requests the next page and stops on a real empty page. IDs are deduplicated within each store. Repeated pages, repeated next links, page ceilings, unavailable routes, invalid rows, published-count mismatches and runtime/budget stops are visible.

`SUMMARY.status` can be `COMPLETED`, `EMPTY`, `PARTIAL`, `FAILED` or `UNCERTAIN`; an individual store can also have `LIMIT_REACHED` or `DUPLICATE_INPUT`. `fullCatalog=true` is reserved for primary-route exhaustion without invalid rows or a known count mismatch. A run reaching an explicitly requested positive product limit is a successful bounded export, not a full catalogue. A collection fallback never claims full-catalogue coverage.

An uncertain default-dataset write or charge stops the whole run immediately. `uncertainDefaultWrites` identifies the product that may already have been persisted or charged; `rowsStored` then counts only acknowledged writes and `rowsStoredCountIsExact=false`. No blind write retry occurs. A copy-dataset failure also stops the run after preserving the acknowledged default row. Copy writes are not atomic with default writes, and an ambiguous copy may already exist: review before manually resuming.

Cloud status alone is not a data-completeness claim. Partial, failed and uncertain runs are failed explicitly after diagnostics are saved. A platform hard kill may interrupt final summary persistence; choose a lower `maxRunSeconds` than the run timeout. Current code does not implement cross-run resume or migration checkpoint recovery.

### Security and access

The optional `datasetId` resource picker declares `READ` and `WRITE` for the selected dataset under Apify limited permissions: metadata lookup requires read access and appending copied products requires write access. Apify expands the run token's scope for the dataset supplied in input; full-account permissions are not required. An omitted or empty `datasetId` triggers no external dataset access. `runId` is only a row label and requests no run access. A storage permission failure is a setup failure before source collection. See [Apify's user-provided storage permissions](https://docs.apify.com/actors/development/permissions#access-user-provided-storages).

Public GET requests and one fixed read-only tokenless Storefront GraphQL POST operation are used. The POST reads only one public product and its variants; it accepts no caller-supplied query, token, customer or admin fields. All POST redirects are rejected. The same verified IP/TLS/proxy transport is used for both methods. HTTP(S) URLs, redirect destinations and DNS answers are checked. Private, loopback, link-local, multicast, fake benchmark and transition addresses are rejected. Numeric-IP connections retain the original Host header and TLS SNI. The same policy also applies to configured proxy endpoints and destination IPs in CONNECT requests. Response bodies and redirect chains have limits. Tokens and cookies are not accepted as store input.

Custom HTTP proxies and Apify Proxy are implemented. HTTPS-scheme proxy servers and SOCKS proxies are an explicit limitation; TLS is not weakened to support them. Proxy setup errors are surfaced, and no implicit proxy purchase or automatic switching occurs. Paid proxy connectivity has only mock coverage in this delivery. Login, password pages and access denials remain boundaries. No cart-add requests, checkout automation, CAPTCHA solving, private Admin API, or authenticated storefront API are used.

### Large variant sets and source limits

Shopify’s Ajax product API documents a maximum 250 variants per response. For catalogues exposing 100 or more variants (including the historical 100 boundary), or reporting a larger explicit count, this Actor verifies and recovers the product with tokenless Storefront API 2026-07 cursor pages of 100. Recovery requires matching product ID/handle and variant parents on every page, unique IDs/cursors, consistent MoneyV2 currencies, an EXACT total count and a terminal page. The hard ceiling is 25 variant pages per product, also subject to the run deadline and transport retries. A failed recovery skips the incomplete product and marks the run partial; it never stores the prefix as a complete price range or exact inventory. The tokenless query does not request quantityAvailable or private inventory. A local live check on September 26, 2026 recovered 280 variants in three pages after both JSON endpoints returned 250. Separate cloud acceptance covered six products, including one with 100 variants; it was not a cloud test of the 280-variant product. Private/draft/wholesale/password-protected inventory, location-level counts, unavailable metafields and headless catalogues without compatible public routes remain outside the accessible source. Source changes during pagination can still alter a snapshot; count matching alone cannot guarantee atomicity.

### API and scheduled exports

Save your input as `input.json`, then start a run with your own Apify token:

```sh
curl --fail-with-body -X POST \
  'https://api.apify.com/v2/acts/oQCaw1otYIwHhIYfm/runs' \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  --data-binary @input.json
```

Read the returned run's default dataset for rows and its key-value store `SUMMARY` record for coverage. Export JSON, CSV or Excel in Apify Console. You can configure an Apify schedule for repeated snapshots; this Actor does not create schedules itself.

Use `(store, productId)` and variant IDs to compare snapshots. `datasetId` appends copies; it does not deduplicate across runs or replace earlier snapshots.

### Support

Use this Actor's **Issues** tab for bugs. Include the run ID, a public store URL and the relevant `SUMMARY` error; never post your Apify token or store credentials. Availability is not an exact stock count. Private inventory and password-protected stores are outside this tool's scope.

### Proxy connection compatibility

On Apify, the SDK can supply a private internal proxy gateway. Only its exact host and port from the platform configuration are accepted as a proxy connection when Apify Proxy is enabled. Store targets and user-supplied proxy URLs remain subject to public-address checks; product connections preserve pinned public IPs and verified TLS.

# Actor input Schema

## `domains` (type: `array`):

Stores to read. Bare hosts and HTTP(S) URLs are accepted; paths are ignored. An empty list uses deathwishcoffee.com. This input does not discover stores.

## `maxProducts` (type: `integer`):

0 reads the complete available catalogue until an empty page, a documented source limit, or a safety/time/budget stop. Positive values cap stored products separately for each store.

## `includeInventoryDetails` (type: `boolean`):

Read public product JSON, Ajax JSON and HTML for source-exposed quantities, barcode, policies, subscriptions, metadata and structured brand. No cart mutation or quantity probing. Unavailable fields are omitted.

## `proxyConfiguration` (type: `object`):

The Console trial and automatic test prefill use Apify RESIDENTIAL proxy, which incurs separate proxy usage costs. If omitted in JSON, proxy remains disabled. Apify proxy settings or user-supplied HTTP CONNECT proxyUrls route requests. Direct cloud requests can receive HTTP 429. HTTPS-scheme and SOCKS proxies are not supported; configuration failures are explicit.

## `datasetId` (type: `string`):

Optional existing dataset receives each stored product in addition to the default dataset. Read/write access is requested only for the selected dataset. Leave unselected or empty to use only the default dataset. Failure to verify it stops setup; append failures stop with partial status.

## `runId` (type: `string`):

Optional 17-character ID copied verbatim onto each product. Omitted from output when empty.

## `locale` (type: `string`):

Explicit locale prefix such as fr or en-gb. Store must expose that route. Domain paths are otherwise ignored. Currency is observed, never converted.

## `pageSize` (type: `integer`):

Pagination batch size, not a total result limit. The same size is used on every page.

## `maxPagesPerStore` (type: `integer`):

Emergency stop. Reaching this limit sets partial status and never claims a complete catalogue.

## `requestTimeoutSecs` (type: `integer`):

Finite timeout per request attempt.

## `maxRetries` (type: `integer`):

Retry transient network failures, 429 and selected 5xx responses. Access denial is not retried or bypassed.

## `requestIntervalSecs` (type: `number`):

Sequential requests with this minimum start interval; no concurrent store crawling.

## `maxRunSeconds` (type: `integer`):

Stops collection and records incomplete progress in SUMMARY. Set below the platform timeout to leave time for final storage.

## Actor input object example

```json
{
  "domains": [
    "deathwishcoffee.com"
  ],
  "maxProducts": 3,
  "includeInventoryDetails": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "locale": "",
  "pageSize": 3,
  "maxPagesPerStore": 10000,
  "requestTimeoutSecs": 25,
  "maxRetries": 2,
  "requestIntervalSecs": 0.5,
  "maxRunSeconds": 90
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "deathwishcoffee.com"
    ],
    "maxProducts": 3,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    },
    "pageSize": 3,
    "maxRunSeconds": 90
};

// Run the Actor and wait for it to finish
const run = await client.actor("peerless_columbine/shopify-products-inventory-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": ["deathwishcoffee.com"],
    "maxProducts": 3,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
    "pageSize": 3,
    "maxRunSeconds": 90,
}

# Run the Actor and wait for it to finish
run = client.actor("peerless_columbine/shopify-products-inventory-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "deathwishcoffee.com"
  ],
  "maxProducts": 3,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "pageSize": 3,
  "maxRunSeconds": 90
}' |
apify call peerless_columbine/shopify-products-inventory-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,peerless_columbine/shopify-products-inventory-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oQCaw1otYIwHhIYfm/builds/OpZf7vuFlIXz4oQQO/openapi.json
