# Walmart Product Scraper — Full Product Pages (`thenetaji/walmart-product-scraper`) Actor

Read Walmart product pages in bulk. Paste one product link or a whole list and each comes back as a structured record of everything the page publishes.

- **URL**: https://apify.com/thenetaji/walmart-product-scraper.md
- **Developed by:** [The Netaji](https://apify.com/thenetaji) (community)
- **Categories:** E-commerce, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.25 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Walmart Product Scraper

The Actor reads Walmart product pages in bulk and saves each as a structured record containing the product name, price and currency, brand, model, star rating and review count, the seller, stock state, the return policy, and the full image set. Products that have been withdrawn produce no dataset record rather than an error.

```json
{
  "product_urls": [
    "https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971",
    "https://www.walmart.com/ip/Keurig-K-Express-Coffee-Maker/87654321"
  ]
}
```

### Accepted input

`product_urls` is required and accepts a list of Walmart product page links in their full `/ip/<name>/<id>` form. Both an absolute link and a bare path are accepted, and a link carrying Walmart's own tracking query is accepted unchanged. Absolute links must be on `walmart.com`; a link on any other host is refused rather than forwarded.

A bare item ID is not accepted, and this is a property of Walmart rather than a limitation of the input parsing. Walmart's anti-bot gate treats the shortened `/ip/<id>` shape differently from the real `/ip/<name>/<id>` one, so a link assembled from an ID alone is rejected upstream. Refusing it here turns what would otherwise be an opaque upstream failure into a legible error naming the entry that caused it.

### Result fields

`item_id` is Walmart's item identifier and `url` is the canonical product page link as the page reports it, which may differ from the link supplied as input.

`name`, `price`, `currency`, `brand`, `model`, and `upc` describe the product. `rating` and `review_count` cover its reviews, `seller_id` and `seller_name` the seller, and `availability` its stock state. `return_policy` states whether the product is returnable, whether returns are free, and the length of the return window. `variants` lists the purchasable variants, `image_url` and `image_info` cover the imagery, `category` is the Walmart category path, and `short_description` is the listing's marketing copy. `product_detail` carries the complete page record for anything not broken out into its own field.

```json
{
  "item_id": "5162907971",
  "url": "https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971",
  "name": "Mainstays Black 12-Cup Drip Coffee Maker",
  "price": 16.88,
  "currency": "USD",
  "brand": "Mainstays",
  "model": "MS8402550614-06",
  "upc": null,
  "rating": 4.5,
  "review_count": 10730,
  "seller_name": "Walmart.com",
  "availability": "In stock",
  "return_policy": { "returnable": true, "freeReturns": true, "returnPolicyText": "Free 90-day returns" }
}
```

This is a trimmed, live-verified result. `upc` is `null` on this product; the field is published but not populated for every listing, and the same is true of `variants` on a product sold in a single configuration.

### The product page is not a larger search row

A product page and a search result describe the same product in different vocabularies, and only around forty field names are common to both. Two differences are worth knowing when combining exports from this Actor and the [Walmart Search Scraper](https://apify.com/thenetaji/walmart-search-scraper).

The page publishes no top-level price. It quotes the figure under its own price information together with a currency, and `price` here is read from there; a search row quotes it at the top level and carries no currency at all. Separately, `brand`, `model`, `upc`, and the return policy exist only on the page — a search row's `brand` is `null` on most listings, which is why enriching a search run is the only way to get those columns without a second pass.

### Behaviour on partial results

Each product is a separate page and therefore a separate request. Three outcomes are handled without stopping the run. An entry that is not a usable Walmart product link is reported and skipped before any request is made, so it costs nothing. A product that has been withdrawn is reported by the source as an empty response rather than as a failure, and no row is saved for it. A request that fails outright is logged and skipped.

A run of ten links in which two are dead therefore finishes normally with eight rows saved. A run whose `product_urls` list is empty, or in which no entry is a usable product link at all, is rejected before any request is made and names the required link form.

Walmart gates product pages more aggressively than search. A request that is turned away is retried, and a product that cannot be read after those retries is reported as a retryable failure rather than saved as a partial row; re-running the same links usually succeeds.

### Getting product links in bulk

The practical constraint on this Actor is that it needs real product links, and Walmart does not publish a list of them. The [Walmart Search Scraper](https://apify.com/thenetaji/walmart-search-scraper) is the usual source: every row it saves carries a `url` in exactly the form required here. Where full product detail is wanted for an entire search rather than a hand-picked list, enabling that Actor's own product detail add-on does the same work in one run and avoids exporting links and re-importing them.

### Related Actors

The [Walmart Search Scraper](https://apify.com/thenetaji/walmart-search-scraper) finds products by keyword and supplies the links this Actor takes. For the same job on other marketplaces, the [eBay Product Scraper](https://apify.com/thenetaji/ebay-product-scraper) reads eBay listings in full and accepts a bare item ID as well as a link, and the [Etsy Search Scraper](https://apify.com/thenetaji/etsy-search-scraper) searches Etsy by keyword.

# Actor input Schema

## `product_urls` (type: `array`):

One or more Walmart product page links, in their full /ip/<name>/<id> form. A bare item ID is not accepted — Walmart rejects links built from one.

## Actor input object example

```json
{
  "product_urls": [
    "https://www.walmart.com/ip/Great-Value-Coffee/12345678"
  ]
}
```

# Actor output Schema

## `dataset` (type: `string`):

All records scraped by this run

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "product_urls": [
        "https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("thenetaji/walmart-product-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "product_urls": ["https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971"] }

# Run the Actor and wait for it to finish
run = client.actor("thenetaji/walmart-product-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "product_urls": [
    "https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971"
  ]
}' |
apify call thenetaji/walmart-product-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thenetaji/walmart-product-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JNZdr1Hc1IDO8ABDz/builds/eo6Zdl4RsBw9HjGnm/openapi.json
