Walmart Product Scraper — Full Product Pages avatar

Walmart Product Scraper — Full Product Pages

Pricing

from $4.25 / 1,000 results

Go to Apify Store
Walmart Product Scraper — Full Product Pages

Walmart Product Scraper — Full Product Pages

Read Walmart product pages in bulk. Paste one product link or a whole list and each comes back as a structured record of everything the page publishes.

Pricing

from $4.25 / 1,000 results

Rating

0.0

(0)

Developer

The Netaji

The Netaji

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 days ago

Last modified

Share

Walmart Product Scraper

The Actor reads Walmart product pages in bulk and saves each as a structured record containing the product name, price and currency, brand, model, star rating and review count, the seller, stock state, the return policy, and the full image set. Products that have been withdrawn produce no dataset record rather than an error.

{
"product_urls": [
"https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971",
"https://www.walmart.com/ip/Keurig-K-Express-Coffee-Maker/87654321"
]
}

Accepted input

product_urls is required and accepts a list of Walmart product page links in their full /ip/<name>/<id> form. Both an absolute link and a bare path are accepted, and a link carrying Walmart's own tracking query is accepted unchanged. Absolute links must be on walmart.com; a link on any other host is refused rather than forwarded.

A bare item ID is not accepted, and this is a property of Walmart rather than a limitation of the input parsing. Walmart's anti-bot gate treats the shortened /ip/<id> shape differently from the real /ip/<name>/<id> one, so a link assembled from an ID alone is rejected upstream. Refusing it here turns what would otherwise be an opaque upstream failure into a legible error naming the entry that caused it.

Result fields

item_id is Walmart's item identifier and url is the canonical product page link as the page reports it, which may differ from the link supplied as input.

name, price, currency, brand, model, and upc describe the product. rating and review_count cover its reviews, seller_id and seller_name the seller, and availability its stock state. return_policy states whether the product is returnable, whether returns are free, and the length of the return window. variants lists the purchasable variants, image_url and image_info cover the imagery, category is the Walmart category path, and short_description is the listing's marketing copy. product_detail carries the complete page record for anything not broken out into its own field.

{
"item_id": "5162907971",
"url": "https://www.walmart.com/ip/Mainstays-Black-12-Cup-Drip-Coffee-Maker/5162907971",
"name": "Mainstays Black 12-Cup Drip Coffee Maker",
"price": 16.88,
"currency": "USD",
"brand": "Mainstays",
"model": "MS8402550614-06",
"upc": null,
"rating": 4.5,
"review_count": 10730,
"seller_name": "Walmart.com",
"availability": "In stock",
"return_policy": { "returnable": true, "freeReturns": true, "returnPolicyText": "Free 90-day returns" }
}

This is a trimmed, live-verified result. upc is null on this product; the field is published but not populated for every listing, and the same is true of variants on a product sold in a single configuration.

The product page is not a larger search row

A product page and a search result describe the same product in different vocabularies, and only around forty field names are common to both. Two differences are worth knowing when combining exports from this Actor and the Walmart Search Scraper.

The page publishes no top-level price. It quotes the figure under its own price information together with a currency, and price here is read from there; a search row quotes it at the top level and carries no currency at all. Separately, brand, model, upc, and the return policy exist only on the page — a search row's brand is null on most listings, which is why enriching a search run is the only way to get those columns without a second pass.

Behaviour on partial results

Each product is a separate page and therefore a separate request. Three outcomes are handled without stopping the run. An entry that is not a usable Walmart product link is reported and skipped before any request is made, so it costs nothing. A product that has been withdrawn is reported by the source as an empty response rather than as a failure, and no row is saved for it. A request that fails outright is logged and skipped.

A run of ten links in which two are dead therefore finishes normally with eight rows saved. A run whose product_urls list is empty, or in which no entry is a usable product link at all, is rejected before any request is made and names the required link form.

Walmart gates product pages more aggressively than search. A request that is turned away is retried, and a product that cannot be read after those retries is reported as a retryable failure rather than saved as a partial row; re-running the same links usually succeeds.

The practical constraint on this Actor is that it needs real product links, and Walmart does not publish a list of them. The Walmart Search Scraper is the usual source: every row it saves carries a url in exactly the form required here. Where full product detail is wanted for an entire search rather than a hand-picked list, enabling that Actor's own product detail add-on does the same work in one run and avoids exporting links and re-importing them.

The Walmart Search Scraper finds products by keyword and supplies the links this Actor takes. For the same job on other marketplaces, the eBay Product Scraper reads eBay listings in full and accepts a bare item ID as well as a link, and the Etsy Search Scraper searches Etsy by keyword.