# Walmart Mexico Product & Price Data Scraper (`w3crawler/walmart-mexico-scraper`) Actor

Extracts publicly embedded product and price data from Walmart Mexico pages.

- **URL**: https://apify.com/w3crawler/walmart-mexico-scraper.md
- **Developed by:** [w3crawler](https://apify.com/w3crawler) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 products & prices

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Walmart Mexico Product & Price Data Scraper do?

Walmart Mexico Product & Price Data Scraper extracts product records from public `walmart.com.mx` home, browse, search, and product pages. It reads JSON-LD, embedded server state, and legacy public product payloads, then normalizes product identity, title, canonical URL, MXN price, previous price, brand, rating, review count, availability, image, category, description, and SKU when those values are publicly exposed.

The Actor uses ordinary HTTPS requests with a transparent user agent. A browser-rendered retry is used only when a successful public response is reachable but empty or not parseable. It does not log in, solve CAPTCHAs, bypass access controls, or fabricate product data.

### Why use Walmart Mexico Product & Price Data Scraper?

Use it for Mexico catalog snapshots, MXN price monitoring, availability checks, product research, and assortment analysis. Each successful row keeps the requested source URL, final response URL, market, rank, transport, extraction method, and capture timestamp so downstream users can audit the result.

### What data can Walmart Mexico Product & Price Data Scraper extract?

| Group       | Fields                                                                                                        |
| ----------- | ------------------------------------------------------------------------------------------------------------- |
| Identity    | `id`, `title`, `url`, `brand`, `sku`, `category`                                                              |
| Commerce    | `price`, `priceText`, `previousPrice`, `currency`, `availability`                                             |
| Content     | `description`, `imageUrl`, `rating`, `reviewCount`                                                            |
| Provenance  | `source`, `sourceUrl`, `finalUrl`, `market`, `rank`, `sourceTransport`, `extractionMethod`, `scrapedAt`       |
| Diagnostics | `recordType`, `status`, `dataAvailable`, `sourceBlocked`, `diagnosticCode`, `diagnosticMessage`, `httpStatus` |

Rows are source-backed. Optional fields are omitted when Walmart does not publish them. Incomplete candidates are counted in `skippedIncompleteCount` instead of being emitted as misleading product rows.

### How to scrape Walmart Mexico products

1. Open the Input tab and enter one or more public HTTPS Walmart Mexico URLs.
2. Set `maxItems` for the global number of distinct products to emit.
3. Set `timeoutMs` or enable ordinary Apify Proxy when public direct access requires it.
4. Start the Actor and inspect the dataset and the `OUTPUT` key-value record.

Only `walmart.com.mx` and its subdomains are accepted. Invalid or missing URLs become explicit diagnostics; the Actor does not silently substitute unrelated pages.

### How much will it cost to scrape Walmart Mexico products?

Apify bills the run according to the account and pricing configuration shown in the Console. Cost depends on the number of URLs, response sizes, optional rendered fallback requests, proxy/container resources, and stored rows. Start with one URL and a small `maxItems`, and keep `timeoutMs` bounded for predictable runs.

### Input

Supported settings:

- `startUrls`: an array of up to 50 objects with a public `https://...walmart.com.mx/...` `url`. The default is the public Walmart Mexico home page.
- `maxItems`: global distinct-product limit from 1 to 200; defaults to 20.
- `timeoutMs`: per-request timeout from 5,000 to 120,000 milliseconds; defaults to 30,000.
- `proxyConfiguration`: optional standard Apify Proxy settings. Proxy use is not an access-control bypass.

#### Search or browse page

```json
{
  "startUrls": [{ "url": "https://www.walmart.com.mx/search?q=teclado" }],
  "maxItems": 3,
  "timeoutMs": 30000
}
```

#### Product page with public proxy fallback

```json
{
  "startUrls": [
    { "url": "https://www.walmart.com.mx/ip/example-product/00489711693180" }
  ],
  "maxItems": 1,
  "timeoutMs": 30000,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

### Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

#### Product row

```json
{
  "recordType": "product",
  "status": "success",
  "dataAvailable": true,
  "source": "walmart_mx_public_search",
  "provenance": {
    "requestedUrl": "https://www.walmart.com.mx/search?q=teclado",
    "responseUrl": "https://www.walmart.com.mx/search?q=teclado",
    "market": "MX",
    "rank": 1,
    "id": "EXAMPLE123",
    "extractionMethod": "embedded_product_payload"
  },
  "sourceTransport": "got_scraping_apify_proxy",
  "extractionMethod": "embedded_product_payload",
  "market": "MX",
  "sourceBlocked": false,
  "sourceUrl": "https://www.walmart.com.mx/search?q=teclado",
  "rank": 1,
  "id": "EXAMPLE123",
  "title": "Example product",
  "url": "https://www.walmart.com.mx/ip/example-product/00489711693180",
  "brand": "Example Brand",
  "price": 299,
  "priceText": "$299.00",
  "previousPrice": 399,
  "currency": "MXN",
  "availability": "Disponible",
  "imageUrl": "https://www.walmart.com.mx/example-image.jpg",
  "scrapedAt": "2026-09-08T12:00:00.000Z"
}
```

#### Diagnostic row

```json
{
  "recordType": "diagnostic",
  "status": "error",
  "dataAvailable": false,
  "source": "walmart_mx_public_search",
  "provenance": {
    "requestedUrl": "https://www.walmart.com.mx/search?q=teclado",
    "responseUrl": "https://www.walmart.com.mx/blocked",
    "market": "MX",
    "httpStatus": 412
  },
  "sourceTransport": "got_scraping_direct",
  "extractionMethod": "diagnostic",
  "market": "MX",
  "sourceBlocked": true,
  "sourceUrl": "https://www.walmart.com.mx/search?q=teclado",
  "diagnosticCode": "ACCESS_BOUNDARY",
  "diagnosticMessage": "Access boundary HTTP 412",
  "httpStatus": 412,
  "scrapedAt": "2026-09-08T12:00:00.000Z"
}
```

The `OUTPUT` key-value record contains run status, source and market, URL and product counts, diagnostics, blocked responses, rendered fallbacks, skipped incomplete candidates, field coverage, and completion time.

### Tips and advanced options

Use one search URL for multiple results and one product URL when you need a detail-oriented check. Keep `maxItems` global in mind when supplying multiple pages. A public HTTP 200 response with no parseable product may receive one rendered retry; a challenge response is preserved as an access diagnostic and is not retried through browser automation. Preserve `provenance` and `sourceTransport` when joining datasets.

### FAQ, support, and responsible use

If a run returns only diagnostics, inspect `diagnosticCode`, `httpStatus`, `sourceBlocked`, `provenance.responseUrl`, and `OUTPUT`. `NO_PUBLIC_PRODUCTS` means the response was reachable but yielded no usable product rows; `ACCESS_BOUNDARY` means a protected response was detected. For support, use the Actor Issues tab or API tab and include the run ID, sanitized input shape, and summary counts.

This Actor is not affiliated with Walmart. It extracts only public Walmart Mexico page data and does not access private user information, authenticate, solve CAPTCHAs, or bypass access controls. Follow Walmart's terms, robots guidance, rate limits, applicable law, and privacy obligations.

### Local verification

Run `npm test`, `npm run check`, `npm run schema`, `apify run --purge --input-file qa-input-280.json`, and `npm run validate`. The checked-in fixtures test parser behavior; live pages can change and a diagnostic is an honest result when public access is unavailable.

# Actor input Schema

## `startUrls` (type: `array`):

Public walmart.com.mx URLs to inspect.

## `maxItems` (type: `integer`):

Maximum product records to emit.

## `timeoutMs` (type: `integer`):

Maximum time for each public page request.

## `proxyConfiguration` (type: `object`):

Optional Apify Proxy. Direct access is used by default.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.walmart.com.mx/"
    }
  ],
  "maxItems": 20,
  "timeoutMs": 30000,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("w3crawler/walmart-mexico-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("w3crawler/walmart-mexico-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call w3crawler/walmart-mexico-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,w3crawler/walmart-mexico-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QjLpQehnVLThkEhra/builds/XRXVRb5FaCsoS5h0q/openapi.json
