# Walmart Intelligence Scraper (`w3crawler/walmart-search-scraper`) Actor

Extracts public Walmart products, prices, ratings, stock, sellers, details, best sellers, deals, category listings, and reviews through bounded source-backed modes.

- **URL**: https://apify.com/w3crawler/walmart-search-scraper.md
- **Developed by:** [w3crawler](https://apify.com/w3crawler) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 product intelligences

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Walmart Search Scraper do?

Walmart Search Scraper collects source-backed public product, category, product-detail, best-seller, deal, and review records from Walmart.com or Walmart.ca. Search mode is the default. The Actor preserves market-specific prices and currencies, seller and availability fields, ratings, review counts, images, source provenance, and structured diagnostics. It does not sign in, solve CAPTCHA, interact with challenges, bypass access controls, or invent product data.

### Why scrape Walmart search results?

Use the Actor for public catalog research, product discovery, price and availability snapshots, category analysis, seller comparisons, and review collection. The bounded modes let you request only the surfaces you need, while the `market` setting keeps Walmart US and Walmart Canada results separate.

### What data can it extract?

The dataset uses one row per source-backed product, detail, category, best-seller, deal, or review record.

| Field                                                                                     | Description                                                          |
| ----------------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| `row_type`, `mode`, `market`, `query`, `position`                                         | Row classification, source mode, market, query, and result position. |
| `item_id`, `us_item_id`, `name`, `brand`, `url`                                           | Public product identity and canonical product URL.                   |
| `price`, `original_price`, `currency`, `rollback`                                         | Source-provided pricing fields when available.                       |
| `in_stock`, `availability_status`                                                         | Source-provided availability fields.                                 |
| `seller_name`, `rating`, `review_count`                                                   | Public seller and review summary values.                             |
| `image_url`, `images`, `category`, `badges`                                               | Public product media and category metadata.                          |
| `model_number`, `upc`, `specifications`, `breadcrumbs`                                    | Additional product-detail fields when exposed.                       |
| `review_id`, `author`, `title`, `text`, `date`, `verified_purchase`, `helpful_count`      | Review fields when review mode is selected.                          |
| `source_website`, `source_url`, `source_transport`, `extraction_method`, `source_blocked` | Source provenance and access status.                                 |

Rows that contain no usable source-backed data are emitted as diagnostic records with `record_type`, `status`, `data_available`, `url`, `error`, `errorCode`, and provenance fields. Missing values remain absent or null.

### How to scrape Walmart

1. Select one or more values in `dataTypes`.
2. Supply `searchQueries` for search mode, `categoryUrls` for category mode, `productIds` or `productUrls` for product/review modes, or `bestsellersDealsCategory` for best-seller/deal modes.
3. Choose `market: "US"` or `market: "CA"` and an optional sort order.
4. Set bounded result, page, review, and concurrency limits.
5. Start the Actor, inspect diagnostics, and export the Dataset.

#### Input example

```json
{
  "dataTypes": ["search"],
  "searchQueries": ["wireless keyboard"],
  "market": "CA",
  "sort": "best_match",
  "maxResultsPerType": 40,
  "maxPages": 5,
  "maxConcurrency": 1,
  "includeProductDetails": false
}
```

`dataTypes` supports `search`, `category`, `product`, `bestsellers`, `deals`, and `reviews`. Search mode defaults to `air fryer` when no query is supplied. URL aliases (`startUrls`, `urls`, and `searchUrls`) and legacy aliases such as `query`, `sortBy`, and `maxItems` remain supported and are validated as official HTTPS Walmart URLs or bounded values.

For category mode, provide one or more official `/browse/` or `/cp/` URLs. For product and review modes, provide numeric item IDs or official `/ip/` URLs. For best-seller and deal modes, provide a category keyword or official category URL.

### Cost and limits

The Actor uses normal Apify compute, browser, proxy, and network resources. Each selected mode is bounded by `maxResultsPerType` (1–500), `maxPages` (1–25), `maxConcurrency` (1–5), and `maxReviewsPerProduct` (1–1,000). Search queries are limited to 20 unique values, and URL arrays have explicit upper bounds. Direct browser access is the default; the optional Apify Proxy configuration is available when authorized and required.

Walmart pages can vary by market and can return access challenges. The Actor records a structured diagnostic instead of using a CAPTCHA solver, third-party reader, login, stealth behavior, or challenge interaction. A diagnostic-only run completes operationally with `COMPLETED_WITH_ERRORS` in `OUTPUT` and is not treated as a successful data scrape.

### Output example

```json
{
  "source": "w3crawler/walmart-search-scraper",
  "source_website": "https://www.walmart.ca/",
  "source_url": "https://www.walmart.ca/en/search?q=wireless+keyboard",
  "query": "wireless keyboard",
  "mode": "search",
  "market": "CA",
  "source_transport": "playwright_direct",
  "extraction_method": "embedded_search_payload",
  "source_blocked": false,
  "scraped_at": "2026-09-08T12:00:00.000Z",
  "row_type": "search_result",
  "position": 1,
  "us_item_id": "6D216IVRTJAC",
  "item_id": "6D216IVRTJAC",
  "name": "Public Walmart product",
  "brand": "Example Brand",
  "price": 84.93,
  "original_price": 99.99,
  "currency": "CAD",
  "rollback": false,
  "rating": 4.5,
  "review_count": 12,
  "in_stock": true,
  "seller_name": "Example Seller",
  "availability_status": "Available",
  "image_url": "https://i5.walmartimages.ca/asr/example.jpeg",
  "images": ["https://i5.walmartimages.ca/asr/example.jpeg"],
  "url": "https://www.walmart.ca/en/ip/Public-Product/6D216IVRTJAC"
}
```

Product-detail rows may also include model, UPC/GTIN, specifications, breadcrumbs, and bounded source data. Review rows use `row_type: "review"` and include review fields when Walmart exposes them. The `OUTPUT` key-value record contains per-mode counts, field counts, pages visited, and diagnostic counts.

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

### Tips for better results

- Set `market` explicitly when collecting Canadian results so prices and source URLs are interpreted correctly.
- Keep `maxResultsPerType` small while testing a new mode.
- Use canonical Walmart URLs and inspect the source page when relying on price or availability.
- Review `source_transport`, `extraction_method`, and `source_blocked` when comparing runs.
- Treat absent fields as source limitations rather than inferred values.
- Use proxy settings only when authorized and permitted for the target market.

### FAQ

#### Why did I receive diagnostics instead of products?

The source page may have returned an access challenge, an empty result, malformed public data, or a non-supported page. Inspect `errorCode`, `url`, and provenance fields in the Dataset.

#### Can I collect Walmart Canada data?

Yes. Set `market` to `CA` and use official Walmart.ca category or product URLs where a URL input is required. Currency and availability are preserved from the source.

#### Does the Actor use a Walmart API or private account?

No. It reads public page content through bounded browser requests and does not log in or access private systems.

### Support, legal notice, and non-affiliation

For support, inspect the run log, Dataset diagnostics, `OUTPUT`, and the Issues tab before reporting a reproducible problem. Use the API tab for programmatic access and the export controls for downstream analysis.

The Actor collects public page data only. Use it for a legitimate purpose, follow Walmart’s terms and applicable law, respect privacy obligations, and honor technical access restrictions. Public product and review information may still have legal or policy constraints in your jurisdiction.

This independent community tool is not sponsored, endorsed, or administered by Walmart Inc. or its affiliates. Walmart and related marks belong to their respective owners.

### Local verification

From this Actor directory, run `npm test`, `npm run lint`, `npm run schema`, and `apify run --purge`. Then run `npm run validate:dataset` against the generated local storage. A qualifying run should contain distinct source-backed rows with zero diagnostics for the selected public mode; challenge-only runs are documented as blocked rather than successful.

# Changelog

This Actor's version history is a separate document: https://apify.com/w3crawler/walmart-search-scraper/changelog.md

# Actor input Schema

## `dataTypes` (type: `array`):

One or more surfaces: search, category, product, bestsellers, deals, or reviews.

## `searchQueries` (type: `array`):

Keywords for search mode.

## `query` (type: `string`):

Compatibility alias for one search query.

## `startUrls` (type: `array`):

Compatibility alias for public Walmart search, category, or product URLs.

## `urls` (type: `array`):

Compatibility alias for startUrls.

## `searchUrls` (type: `array`):

Compatibility alias for public Walmart search URLs.

## `categoryUrls` (type: `array`):

Public https://www.walmart.com/browse/... or https://www.walmart.ca/en/... category URLs.

## `productIds` (type: `array`):

Walmart item IDs for product or review mode.

## `productUrls` (type: `array`):

Public https://www.walmart.com/ip/... or https://www.walmart.ca/en/ip/... product URLs.

## `bestsellersDealsCategory` (type: `string`):

Category keyword, or a public Walmart browse URL, for best-seller and deals modes.

## `sort` (type: `string`):

Search or list ordering.

## `sortBy` (type: `string`):

Compatibility alias for sort.

## `market` (type: `string`):

Public Walmart market to query. The default preserves US behavior; choose CA for Walmart Canada prices and availability.

## `maxResultsPerType` (type: `integer`):

Bounded maximum rows for each selected data type.

## `maxItems` (type: `integer`):

Compatibility alias for maxResultsPerType.

## `maxReviewsPerProduct` (type: `integer`):

Maximum review rows collected for each product page.

## `includeProductDetails` (type: `boolean`):

Queue bounded product detail pages for listing rows.

## `maxPages` (type: `integer`):

Maximum public search pages to inspect.

## `maxConcurrency` (type: `integer`):

Number of browser pages processed concurrently.

## `proxyConfiguration` (type: `object`):

Optional authorized Apify Proxy configuration. No challenge bypass is attempted.

## Actor input object example

```json
{
  "dataTypes": [
    "search"
  ],
  "searchQueries": [
    "air fryer"
  ],
  "sort": "best_match",
  "market": "US",
  "maxResultsPerType": 40,
  "maxReviewsPerProduct": 100,
  "includeProductDetails": false,
  "maxPages": 5,
  "maxConcurrency": 1
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `keyValueStore` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("w3crawler/walmart-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("w3crawler/walmart-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call w3crawler/walmart-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,w3crawler/walmart-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OKDcLFavc2KVQjnr3/builds/H97CRzEX8NiyQHW1S/openapi.json
