Taobao Product Scraper
Under maintenancePricing
$20.00/month + usage
Taobao Product Scraper
Under maintenanceExtract detailed product data from Taobao, including title, price, reviews, specifications, images, options, and categories. Perfect for market research, e-commerce insights, and data-driven decision-making
Pricing
$20.00/month + usage
Rating
0.0
(0)
Developer
Pizani
Maintained by CommunityActor stats
6
Bookmarked
220
Total users
6
Monthly active users
2 days ago
Last modified
Categories
Share
🛒 Taobao Product Scraper
This Actor extracts product data from Taobao and returns a ready-to-use JSON — in English, with real prices. No browser is used, so runs are fast and cheap.
✨ What it does
- Accepts a Taobao product_url (or a bare item ID).
- Extracts title (English + original Chinese), spec sheet, price (USD + CNY), stock, images, description images and the full SKU matrix with per-variant price and stock.
- Collects store data (name, ID and link).
- Pushes the result to the Dataset (Output tab).
🎯 What makes it different
| Typical Taobao scrapers | This Actor | |
|---|---|---|
| Title | Chinese only | English + titleCN |
| Spec sheet | Chinese only | English (Brand, Movement type, Case material, …) |
| Prices | CNY only | USD + CNY, converted at the live rate |
| Variants | raw property IDs | readable names: Pink / 128 Gb / Package 1 |
Prices are converted from the real seller price in CNY using the live exchange rate — they are not a shopping-agent's marked-up price.
🔧 Input
{"product_url": "https://item.taobao.com/item.htm?id=692815138312"}
| Field | Required | Description |
|---|---|---|
product_url | ✅ | Product URL or bare item ID. |
🧾 Example Output
{"sellerInfo": {"shopTitle": "A电遇体育服饰","shopID": "339788445","shopLink": "https://shop339788445.taobao.com","sellerId": null},"productInfo": {"urlproduct": "https://item.taobao.com/item.htm?id=692815138312","itemId": "692815138312","title": "Retro Brazilian Sports Fan Jersey Round Neck Lapel Polo Football Uniform","titleCN": "复古巴西运动球迷球衣圆领翻领Polo杉半拉长拉足球服热身训练服","price": "11.88","priceCurrency": "USD","priceCNY": "68.00","originalPriceCNY": null,"imgList": ["https://img.alicdn.com/...png", "... 4 more ..."],"descriptionImages": ["https://img.alicdn.com/...jpg", "... 10 more ..."],"totalStock": 20509,"minOrderQuantity": 1,"sevenDayReturn": false,"scrapedAt": "2026-08-12T23:01:32+00:00"}}
🧩 Field Reference
sellerInfo.shopTitle/shopID/shopLink/sellerId: Store name, numeric ID, URL and seller account.productInfo.itemId/urlproduct: Numeric Taobao item ID and its canonical URL.productInfo.title/titleCN: English title / original Chinese title.productInfo.price/priceCurrency/priceCNY: USD price / currency (alwaysUSD) / real CNY price.productInfo.originalPriceCNY: Crossed-out price —nullunless there is a real discount.productInfo.imgList/descriptionImages: Gallery images / images inside the description (tracking pixels and duplicates removed).productInfo.totalStock/minOrderQuantity/sevenDayReturn: Stock, MOQ and 7-day return policy.productInfo.atributtes: Spec sheet as key/value, translated. Colour/size axes are not repeated here — they are inoptions.productInfo.options[]: Every SKU —name,imgUrl,price,priceCNY,stock,skuId.productInfo.scrapedAt: UTC timestamp of the run.
🛡️ Reliability
Product data is resolved through multiple independent channels in parallel and merged into a single record. This matters more than it sounds: individual data channels refuse certain products outright — branded listings in particular — so a single-channel scraper simply fails on them. In testing, products that one channel rejected were still returned complete.
A product that no channel can resolve is reported as an explicit error rather than as an empty row, so you never pay for a blank result.
🛠️ How to use via API
import osimport timeimport jsonimport requestsAPI_TOKEN = "YOUR_APIFY_API_KEY"ACTOR_ID = "pizani~taobao-product-scraper"PRODUCT_URL = "https://item.taobao.com/item.htm?id=692815138312"output_dir = "./scraped_results"os.makedirs(output_dir, exist_ok=True)output_file = os.path.join(output_dir, "product_data.json")run_resp = requests.post(f"https://api.apify.com/v2/acts/{ACTOR_ID}/runs?token={API_TOKEN}",json={"product_url": PRODUCT_URL})run_data = run_resp.json()run_id = run_data.get("data", {}).get("id")if not run_id:raise Exception(f"Failed to start Actor: {run_data}")status_url = f"https://api.apify.com/v2/actor-runs/{run_id}?token={API_TOKEN}"while True:status_resp = requests.get(status_url).json()status = status_resp.get("data", {}).get("status")if status in {"SUCCEEDED", "FAILED", "ABORTED", "TIMED-OUT"}:breaktime.sleep(5)items_url = ("https://api.apify.com/v2/datasets/"f"{status_resp['data']['defaultDatasetId']}"f"/items?clean=true&format=json&token={API_TOKEN}")result = requests.get(items_url).json()with open(output_file, "w", encoding="utf-8") as f:json.dump(result, f, ensure_ascii=False, indent=4)
ℹ️ Notes
- Fields can vary depending on the product: not every listing has a spec sheet, a discount or SKU images.
priceCNYis the seller's actual price; converted prices follow the live exchange rate at run time.