# Walmart Product Details Scraper (`w3crawler/walmart-products-scraper`) Actor

Extracts public product details from Walmart product pages.

- **URL**: https://apify.com/w3crawler/walmart-products-scraper.md
- **Developed by:** [w3crawler](https://apify.com/w3crawler) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 products

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Walmart Products Scraper do?

Walmart Products Scraper extracts rich, source-backed product records from public Walmart.com and Walmart.ca product pages. It reads public JSON-LD, embedded page state, and rendered page fields, then emits one normalized product record per requested product URL when the page publishes the required detail fields.

The Actor does not log in, use credentials, solve CAPTCHAs, bypass access controls, or fabricate values. A challenge or incomplete page becomes a structured diagnostic instead of being presented as a product.

### Why use Walmart Products Scraper?

Use it for catalog enrichment, price monitoring, seller comparison, product-page audits, and variant discovery. Each product row preserves the requested URL, final URL, market, rank, transport, extraction method, and timestamp. Apify provides dataset exports, API access, run history, scheduling, and optional ordinary proxy configuration.

### What data does Walmart Products Scraper provide?

| Group       | Fields                                                                                                                                |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| Identity    | `productId`, `name`, `url`, `brand`, `manufacturer`, `sku`, `upc`                                                                     |
| Commerce    | `price`, `originalPrice`, `unitPrice`, `currency`, `availability`, `sellerName`                                                       |
| Detail      | `description`, `images`, `shipping`, `returnPolicy`, `category`, `breadcrumbs`, `features`, `specifications`, `fulfillment`, `badges` |
| Variants    | `variants` when the public page exposes them                                                                                          |
| Provenance  | `source`, `sourceUrl`, `finalUrl`, `market`, `rank`, `sourceTransport`, `extractionMethod`, `scrapedAt`, `httpStatus`                 |
| Diagnostics | `recordType`, `status`, `dataAvailable`, `sourceBlocked`, `diagnosticCode`, `diagnosticMessage`                                       |

Product records are emitted only when price, availability, images, seller, fulfillment, features, and specifications are present. Optional values are omitted when the source does not publish them. Fulfillment and variants are normalized as arrays of source-backed objects.

### How to scrape Walmart products

1. Add one or more public Walmart product URLs under `startUrls`.
2. Set `maxItems` and `maxConcurrency` within their documented limits.
3. Add an authorized Apify Proxy configuration only when your permitted access requires it.
4. Start the Actor and inspect product and diagnostic records together. A diagnostic explains a challenge, request failure, missing public product payload, or incomplete detail page.

The Actor accepts Walmart.com and Walmart.ca HTTPS URLs whose path contains `/ip/`. It processes at most 50 URL entries and at most 50 product pages per run. Local `fixtureFile` input is intended only for deterministic development checks and must point inside this Actor's `fixtures/` directory.

### How much does it cost?

Apify compute and any configured proxy traffic are billed according to your Apify plan. Cost is generally driven by browser time, memory, URL count, and proxy usage. Keep concurrency low and request only the product pages you need; use a fixture locally without network traffic when testing parser changes.

### Input

The default input shape is:

```json
{
  "startUrls": [
    {
      "url": "https://www.walmart.ca/en/ip/DarkBeacon-Flux-68-HE-Magnetic-Switch-Gaming-Keyboard-Adjustable-Actuation-RT/1YTYV6ZNUNUR"
    }
  ],
  "maxItems": 1,
  "maxConcurrency": 1
}
```

Input fields:

| Field                | Type    | Required | Description                                                                                                                  |
| -------------------- | ------- | -------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `startUrls`          | array   | yes      | One to 50 official Walmart.com or Walmart.ca HTTPS `/ip/` product URLs, as strings or `{ "url": "..." }` objects.            |
| `maxItems`           | integer | no       | Number of unique product pages to process; default `1`, range `1–50`.                                                        |
| `maxConcurrency`     | integer | no       | Browser concurrency; default `1`, range `1–3`.                                                                               |
| `proxyConfiguration` | object  | no       | Authorized Apify Proxy settings. `useApifyProxy` must be boolean; optional `proxyUrls` must contain at most 20 HTTP(S) URLs. |
| `fixtureFile`        | string  | no       | Local-only HTML fixture path under `fixtures/`; not a substitute for Cloud source evidence.                                  |

For example, an authorized proxy run can use:

```json
{
  "startUrls": [
    {
      "url": "https://www.walmart.ca/en/ip/DarkBeacon-Flux-68-HE-Magnetic-Switch-Gaming-Keyboard-Adjustable-Actuation-RT/1YTYV6ZNUNUR"
    }
  ],
  "maxItems": 1,
  "maxConcurrency": 1,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

The runtime validator rejects missing or malformed URLs, unsupported hosts or URL schemes, credentials embedded in product URLs, numeric strings for numeric options, invalid ranges, malformed proxy settings, and fixture paths outside `fixtures/`.

### Output

The default dataset contains product and diagnostic records. A representative product record is:

```json
{
  "recordType": "product",
  "status": "success",
  "dataAvailable": true,
  "source": "walmart_public_product_detail",
  "sourceTransport": "playwright_apify_proxy",
  "extractionMethod": "embedded_product_payload_with_idml",
  "market": "CA",
  "sourceBlocked": false,
  "sourceUrl": "https://www.walmart.ca/en/ip/example-product/1YTYV6ZNUNUR",
  "finalUrl": "https://www.walmart.ca/en/ip/example-product/1YTYV6ZNUNUR",
  "rank": 1,
  "productId": "1YTYV6ZNUNUR",
  "name": "Example product",
  "url": "https://www.walmart.ca/en/ip/example-product/1YTYV6ZNUNUR",
  "brand": "Example Brand",
  "price": 49.99,
  "originalPrice": 68.99,
  "currency": "CAD",
  "availability": "In stock",
  "sellerName": "Example Seller",
  "images": ["https://i5.walmartimages.ca/asr/example.jpg"],
  "features": ["Example feature: value"],
  "specifications": { "Manufacturer Part Number": "EXAMPLE-1" },
  "fulfillment": [{ "fulfillment": "DELIVERY" }],
  "variants": [{ "name": "Example variant" }],
  "scrapedAt": "2026-09-08T12:00:00.000Z"
}
```

A diagnostic record has `recordType: "diagnostic"`, `status: "failed"`, `dataAvailable: false`, `sourceBlocked` as a boolean, and a `diagnosticCode` such as `ACCESS_BOUNDARY`, `REQUEST_FAILED`, `NO_PUBLIC_PRODUCT`, or `INCOMPLETE_PRODUCT_DETAIL`. Diagnostics identify why a page did not produce a complete product and are never counted as products.

`OUTPUT` in the default key-value store contains `status`, market, requested and failed pages, product and diagnostic counts, blocked count, skipped incomplete count, field coverage, and completion time.

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

### Tips and advanced options

Use a direct product URL when you need the richest fields. Keep `maxConcurrency` at `1` when diagnosing access boundaries. If the run contains diagnostics, inspect `diagnosticCode`, `diagnosticMessage`, `httpStatus`, `sourceBlocked`, and the source URLs before retrying. A page may be reachable while still omitting fulfillment or specifications; the Actor reports that as `INCOMPLETE_PRODUCT_DETAIL` rather than filling the gap.

### FAQ, support, and responsible use

#### Why did I receive a diagnostic instead of a product?

Walmart may return a public page without a complete embedded product payload, or it may return a human-verification or other access-boundary response. The Actor records the observable result and does not attempt CAPTCHA solving or access-control bypass.

#### Does the Actor guarantee price or availability?

No. These values are time-sensitive source observations, not purchase guarantees. Use Walmart's current product page for decisions that require current price, inventory, delivery, or policy information.

#### How do I request support?

Use the Actor Issues or API tab and include the run ID, sanitized input shape, record counts, and diagnostic code. Do not submit cookies, credentials, private headers, payment information, or personal data.

This Actor is an independent community tool and is not affiliated with, endorsed by, or sponsored by Walmart. Use it only for public, permitted, lawful collection and respect Walmart's terms, robots guidance, rate limits, and applicable law.

### Local verification

Run `npm test`, `npm run check`, `npm run schema`, and `apify run --purge --input-file qa-input-282-fixture.json`, then run `npm run validate`. The fixture is deterministic; a real Cloud run with an authorized input is required to assess current source availability.

# Actor input Schema

## `startUrls` (type: `array`):

Public HTTPS Walmart.com or Walmart.ca product-page URLs. The URL path must contain /ip/ and a product identifier.

## `maxItems` (type: `integer`):

Maximum product pages to process.

## `maxConcurrency` (type: `integer`):

Keep this low to reduce challenge risk.

## `proxyConfiguration` (type: `object`):

Optional Apify Proxy configuration. Walmart may require an authorized proxy.

## `fixtureFile` (type: `string`):

Optional deterministic fixture for local validation. The path must resolve inside this Actor's fixtures/ directory.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.walmart.ca/en/ip/DarkBeacon-Flux-68-HE-Magnetic-Switch-Gaming-Keyboard-Adjustable-Actuation-RT/1YTYV6ZNUNUR"
    }
  ],
  "maxItems": 1,
  "maxConcurrency": 1,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Default dataset containing Walmart product details.

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("w3crawler/walmart-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("w3crawler/walmart-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call w3crawler/walmart-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,w3crawler/walmart-products-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jlD1YArenxphpIDAj/builds/KC3HVgx8v8Wpf2Srk/openapi.json
