# Taobao Scraper (`mlg14/taobao-scraper`) Actor

Fetch public Taobao and Tmall product details by item URL or numeric ID, including listed prices, images, shop details, location, and rating summaries when available.

- **URL**: https://apify.com/mlg14/taobao-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.30 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Taobao Scraper

Collect Taobao and Tmall product details from direct product URLs or numeric item IDs. Each successful request produces one product record with its title, listed price, gallery images, category, public shop details, and rating summary when available. The output can be downloaded as JSON, CSV, or a spreadsheet from the run's dataset.

**Current access status:** The public item-list route returned a real product record in the deployed validation run. That run produced one row for one current product ID. Direct desktop product pages and the mobile detail endpoints still encounter login or security challenges. This actor remains unpublished; test representative IDs before depending on a larger batch because older items can redirect to a generic homepage.

### What data can you extract from Taobao?

A successful row represents one product. The public server-rendered item route supplies basic product details and currently omits the SKU matrix, stock, and full specifications, so those fields are empty in rows from that source. The actor retains variant fields for the mobile detail fallback if it becomes accessible. Values absent from a source response are `null`, empty arrays, or empty objects as appropriate.

| Field | Description | Example or shape |
| --- | --- | --- |
| `input` | The URL or ID originally supplied | `730344766230` |
| `itemId` | Numeric product identifier | `730344766230` |
| `url` | Canonical direct product link | `https://item.taobao.com/item.htm?id=730344766230` |
| `isTmall` | Source flag indicating a Tmall listing | `true` or `false` |
| `title` | Product title as supplied by the detail response | Chinese text |
| `priceCurrency` | Currency assigned to listed product prices | `CNY` |
| `price` | Current price text parsed into a decimal string | `49.90` |
| `originalPrice` | Previous or list price when supplied | `49.90` |
| `priceVerified` | Whether the displayed amount passed independent verification | `false` |
| `hasDiscount` | Whether current and original prices differ | `true` or `false` |
| `brandName` | Brand found among the public specifications | String or `null` |
| `categoryId` | Category identifier from the source | String or `null` |
| `stock` | Product-level quantity when exposed | Integer or `null` |
| `sales` | Sales count as returned by the source | Text, integer, or `null` |
| `skuCount` | Count of parsed SKU combinations | Nonnegative integer |
| `skus` | Variant ID, price, original price, quantity, and option IDs and names | Array of objects |
| `attributes` | Specification name and value pairs | Array of objects |
| `propsList` | Readable choices keyed by option ID | Object |
| `propsImages` | Swatch images keyed by option ID | Object |
| `mainPictureUrl` | First gallery image | URL or `null` |
| `pictures` | Product gallery | Array of URLs |
| `descriptionImages` | Images extracted from a description | Currently empty |
| `descriptionHtml` | Description markup | Currently `null` |
| `shop` | Shop ID, seller ID, name, and URL | Object |
| `location` | Public seller location when available | String or `null` |
| `reviews` | Buyer reviews attached to the product | Currently empty |
| `reviewsFetched` | Whether buyer review collection succeeded | Currently `false` |
| `goodReviewPercent` | Positive rating share displayed publicly | `100%` or `null` |
| `freeShipping` | Whether the public page marks shipping as free | `false` or `null` |
| `detailLevel` | Completeness label for the supported detail response | `partial` |
| `scrapedAt` | UTC time when a successful row was prepared | ISO timestamp |

Prices are strings so that decimal display values are preserved. The actor does not convert CNY into another currency. `priceVerified` stays false because a returned detail value alone is not enough to confirm a live checkout price. `hasDiscount` compares only the two parsed price strings; it does not account for coupons, shipping, customer level, or promotions that apply later. `stock` and `sales` may represent values exposed for a particular region or account state. They should be checked at the product page before purchasing or making a pricing decision.

`skus` is empty for the working public item-list route. If the mobile detail fallback succeeds in a future run, each SKU may include `skuId`, `price`, `originalPrice`, `quantity`, `propsIds`, and `propsNames`. The option IDs can then be joined to `propsList`; `propsImages` may hold swatches. The gallery is in `pictures`, with its first entry repeated as `mainPictureUrl` for tabular exports. Array and object fields retain more structure in JSON than in CSV.

### How to scrape Taobao product details

1. Copy a direct product URL containing an `id` query parameter, or copy the numeric product ID. Direct Taobao and Tmall item links are accepted. A category, search, shop, or short redirect link is not a product ID and is skipped.
2. Add one or more values to `items`. The actor processes distinct IDs in input order. Repeating the same item within a run does not create duplicate rows.
3. Set `maxItems` if the run should stop after a certain number of successful product rows. This cap is on rows delivered, not on URLs attempted. An invalid or blocked URL does not consume the cap.
4. Run the actor and review both the dataset and run log. A successful run process with an empty dataset means no product detail passed validation; it does not mean the requested products were captured.
5. Download the dataset in the format needed for your workflow. Use JSON when variant option mappings or shop objects need to remain structured.

The actor first requests the public server-rendered item-list page at `/list/item/wap/<id>.htm`. Its embedded data must report success, contain a title, and match the requested item ID. If that page lacks usable product data, the actor tries the two mobile detail JSON routes. Login pages, security interstitials, generic homepages, and denied JSON responses do not become product records. This prevents a plausible but false row containing only the supplied ID and URL.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `items` | Array of strings | Required | Direct product URLs with an `id` parameter or numeric IDs; up to 1,000 entries in the form. |
| `maxItems` | Integer | `0` | Maximum successful product rows across the run; zero adds no row cap. |
| `proxyConfiguration` | Object | Platform proxy | Connection settings for product requests. |

A single-product input is useful for checking access before attempting a batch:

```json
{
  "items": ["https://item.taobao.com/item.htm?id=730344766230"],
  "maxItems": 1
}
```

The same validated product can be submitted by numeric ID:

```json
{
  "items": ["730344766230"],
  "maxItems": 1
}
```

The input parser uses the `id` or `itemId` query parameter. It accepts numeric IDs of eight to sixteen digits. It does not follow share links to discover an ID. A malformed input is skipped, and remaining inputs continue. Only domains under Taobao or Tmall are accepted for URLs. A numeric ID can be entered without a URL and is converted to the canonical Taobao item link in output.

### Output example

This is the complete product row from the successful one-item validation run on 2026-09-27. Empty variant and description fields are genuine output from this public route.

```json
{
  "input": "https://item.taobao.com/item.htm?id=730344766230",
  "itemId": "730344766230",
  "url": "https://item.taobao.com/item.htm?id=730344766230",
  "isTmall": false,
  "title": "百亿 农夫山泉纯净水...",
  "priceCurrency": "CNY",
  "price": "49.90",
  "originalPrice": "49.90",
  "priceVerified": false,
  "hasDiscount": false,
  "brandName": null,
  "categoryId": "50008916",
  "stock": null,
  "sales": null,
  "skuCount": 0,
  "skus": [],
  "attributes": [],
  "propsList": {},
  "propsImages": {},
  "mainPictureUrl": "https://img.alicdn.com/imgextra/i3/393647717/O1CN01MpNlgy26sRW0tsXf1_!!393647717.jpg",
  "pictures": [
    "https://img.alicdn.com/imgextra/i3/393647717/O1CN01MpNlgy26sRW0tsXf1_!!393647717.jpg"
  ],
  "descriptionImages": [],
  "descriptionHtml": null,
  "shop": {
    "shopId": "63978904",
    "sellerId": "393647717",
    "shopName": "优送网",
    "shopUrl": "https://shop.m.taobao.com/shop/shop_index.htm?user_id=393647717&item_id=730344766230"
  },
  "location": "上海",
  "reviews": [],
  "reviewsFetched": false,
  "goodReviewPercent": "100%",
  "freeShipping": false,
  "detailLevel": "partial",
  "scrapedAt": "2026-09-27T03:50:25.747149+00:00"
}
```

### Use cases

- **Catalog comparison:** Keep product titles, listed prices, galleries, and public shop identifiers together for side-by-side review. Compare only records collected near the same time because listings can change.
- **Product discovery follow-up:** Turn a list of discovered item IDs into product titles, visible prices, gallery images, and shop links. Check the returned row count because some IDs no longer have a public item-list page.
- **Listing audits:** Check whether a product exposes a gallery, category, shop, and public rating summary. Empty variant and specification fields reflect the limited data on the working route.
- **Price monitoring:** Schedule repeat runs for a fixed list of IDs and compare the latest returned price with earlier successful rows. Since price verification is not available, confirm any significant movement manually.
- **Shop catalog research:** Associate public product listings with public shop identifiers and links. The actor does not collect seller contact details or private account information.

These workflows require successful anonymous product responses. The validation run yielded one row for a current item, while older sample IDs redirected away from their item-list routes. Start with several representative current items and inspect the dataset before planning larger jobs or setting a schedule. A zero-row run is not evidence that every product was deleted.

### How much does it cost to scrape Taobao?

The configured result price is **$0.30 per 1,000 delivered product rows**. Charging is tied to rows pushed to the dataset, subject to the platform's displayed terms and any account-specific usage. Failed or blocked product attempts do not create a dataset row. The current validation run delivered one row, but a single-item run does not establish the efficiency of a large batch.

At the configured result rate, 100 delivered rows correspond to $0.03 in result charges, 1,000 rows correspond to $0.30, and 5,000 rows correspond to $1.50. Those examples describe result charges only and assume access is working. The number of attempted URLs can exceed the number of delivered rows if inputs are invalid, listings are unavailable, or product requests are challenged. Check the run's billing details for the actual amount before relying on an estimate.

### Tips for best results

Use direct product links copied from a product page, with the numeric `id` visible in the query string. Test a current item before submitting a long list. Keep one ID per input entry; multiple URLs pointing to the same ID are deduplicated. Save the original input list so that skipped IDs can be compared with the run log. For structured shop and gallery analysis, export JSON; spreadsheet exports may serialize nested objects as text.

Use `maxItems` when a downstream system expects a fixed maximum. Because it counts successful rows, it cannot guarantee that exactly that many products will appear. A run stops when the cap is reached or when the input list ends. If a data pipeline requires a complete batch, compare the number of distinct valid IDs submitted against the number of rows delivered and investigate the difference.

The public item-list route and mobile detail endpoints can show region-sensitive prices, promotions, logistics, and availability. The working row came from the Chinese-language item-list page; the fallback asks the international mobile detail view for a United States region value. The actor does not expose a region selector because alternate regions have not been validated. Check the price on the target page before using it to place an order.

### Limits

The principal current limit is **partial data and uneven coverage**. The public item-list route supplied one current product's title, price, gallery, category, location, rating summary, and shop. The same route redirected to a generic homepage for two older sample IDs. Its successful response did not include SKU combinations, full attributes, or stock. Desktop product pages returned a login redirect; a mobile page returned an app shell; the newer JSON detail route returned a security denial and the older route returned a challenge script. Residential probing did not unlock them.

The actor does not log in, accept account cookies, solve account-specific verification, or access private data. It does not collect buyer reviews. `reviews` is therefore an empty array and `reviewsFetched` is false in output from the current implementation, even though the public page may show a rating summary or a sample comment. Description HTML and description image parsing are also not active. The fields remain in the dataset shape to make their status clear, not to imply that they are populated.

A direct product ID represents at most one output row, so a single-item validation expects one row rather than thirty. For successful responses, a missing title causes the row to be rejected. A price may still be `null` if the product detail does not expose a usable amount. A short link cannot be resolved by this version, and an unsupported input is skipped rather than submitted to the product endpoint.

### FAQ

#### Why did the run finish with zero rows?

The process can finish normally even when no product response contains usable detail. Review the log for a redirect, security denial, challenge, or unsupported input. Some older IDs redirected from the public item-list route to a generic homepage; those responses were rejected. Zero rows should be treated as a failed data collection outcome for that input, even though the process itself completed.

#### Can I use a numeric ID instead of a link?

Yes. Enter the eight-to-sixteen-digit product ID as a string in `items`. The actor builds the public item-list URL and emits the canonical direct item link if it receives a valid product response. A numeric ID by itself is not enough to create a product row.

#### Are Tmall products supported?

Direct Tmall item URLs with an `id` parameter are accepted and routed by product ID. The source may mark a listing as a Tmall item in its public response. The tested Tmall sample reached a login flow and its item-list route did not yield product data, so Tmall output remains unverified.

#### Why is the price field empty or unverified?

A product can omit a current price, or pricing can depend on selected variant, region, promotion, or account state. The actor does not infer a price from the URL and does not claim checkout verification. A returned price remains an indicative listed value, with `priceVerified` set to false.

#### Does this collect reviews or buyer information?

No. Anonymous review retrieval was not established. The actor does not request private accounts or buyer profiles. The `reviews` field is empty and `reviewsFetched` is false under the present implementation. Do not use it for review analysis.

#### Do I need to configure a proxy?

The default connection settings are applied automatically. The working item-list page responded without a residential connection. Residential probing of direct pages and mobile detail endpoints still reached access controls. Changing the proxy alone is not a demonstrated way to recover those additional fields.

#### Can I schedule runs or export to a spreadsheet?

The platform supports scheduling and dataset exports, including spreadsheet formats. Keep the source IDs in your own input list and compare row counts after each run so that a redirect or new access failure is visible. Verify several representative products before scheduling a large batch.

#### Is collecting this data allowed?

Use only data you are permitted to collect and process. Check the target site's terms, applicable law, and any obligations relating to personal data. This actor is designed around public product details and does not provide a login or collect private account information.

### Integrations

A successful dataset can be downloaded, requested through the platform API, or sent to a workflow through a run completion webhook. Structured JSON preserves `shop`, `pictures`, and any future variant data. An integration should reject an empty dataset when it expected products, record the run ID, and alert the operator before replacing an older product catalog with zero rows. The mixed behavior of current and older sample IDs makes this check essential.

### Support

Open an issue on the Issues tab with the input JSON, run ID, and a short description of the expected product. Include a direct link when an item unexpectedly redirects away from its public item-list page.

# Actor input Schema

## `items` (type: `array`):

Direct Taobao or Tmall product URLs with an id parameter, or numeric product IDs. Each successful product produces one row. Short share links are not supported.

## `maxItems` (type: `integer`):

Stop after this many successful product rows. Zero means no additional cap.

## `proxyConfiguration` (type: `object`):

Connection settings for product requests.

## Actor input object example

```json
{
  "items": [
    "https://item.taobao.com/item.htm?id=730344766230"
  ],
  "maxItems": 0,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `products` (type: `string`):

Successful product details in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "items": [
        "https://item.taobao.com/item.htm?id=730344766230"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/taobao-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "items": ["https://item.taobao.com/item.htm?id=730344766230"] }

# Run the Actor and wait for it to finish
run = client.actor("mlg14/taobao-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "items": [
    "https://item.taobao.com/item.htm?id=730344766230"
  ]
}' |
apify call mlg14/taobao-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/taobao-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sDJQQ0fY093GzOe2N/builds/GQHR6oRezncbzm7Kb/openapi.json
