# Alibaba Category Scraper (`w3crawler/alibaba-category-scraper`) Actor

Extract comprehensive product data from Alibaba's marketplace using keyword searches. This Apify Actor scrapes details like prices, supplier ratings, images, and certifications, delivering structured JSON output for easy analysis....

- **URL**: https://apify.com/w3crawler/alibaba-category-scraper.md
- **Developed by:** [w3crawler](https://apify.com/w3crawler) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 category products

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### Alibaba Category Scraper

Browser-first extraction for public [Alibaba.com](https://www.alibaba.com/) category, search, catalog, and product-detail pages. The Actor collects publicly exposed product and supplier facts from the rendered page, embedded offer data, SSR cards, and Product/ItemList JSON-LD. It supports multiple sources, bounded pagination, deduplication, filters, sorting, standard Apify Proxy routing, and fail-closed diagnostics.

The Actor is not affiliated with Alibaba Group. It requests only public pages, does not sign in, solve CAPTCHAs, evade WAFs, rotate identities to bypass access controls, or call private APIs. When a public response is challenged or unavailable, it stops at that boundary and reports a bounded diagnostic instead of fabricating a product. Availability can vary by region, session, and Cloud transport.

### Why use this Actor

Use it to turn public Alibaba category or search pages into analysis-ready product rows for supplier discovery, preliminary sourcing research, price/MOQ comparisons, catalog monitoring, and market reconnaissance. Multiple input URLs can be crawled in one bounded run, and filters can narrow the result set before the global `maxItems` limit is applied.

### What is extracted

When the public response exposes the value, product rows can include:

- identity: `url`, `productUrl`, `productId`, `productName`, `title`
- pricing: `price`, `priceCurrency`, `priceFormatted`, `priceMin`, `priceMax`, `priceUnit`, `discount`, `promotionPrice`
- order terms: `minOrder`, `minOrderUnit`, `minOrderFormatted`
- supplier: `supplierName`, `supplierProfileUrl`, `supplierCountry`, `supplierLocation`, `supplierYears`, `supplierRating`, `supplierRatingValue`, `supplierId`
- engagement: `ratingsCount`, `ratingsCountValue`, `soldCount`, `soldCountValue`, `transactionCount`, `transactionCountValue`
- catalog facts: `category`, `brand`, `availability`, `productScore`, `supplierServiceScore`, `shippingScore`
- media and merchandising: `images`, `thumbnail`, `imageCount`, `badges`, `certifications`, `sellingPoints`, `isAd`, `isCertified`
- provenance: `sourceUrl`, `pageNumber`, `position`, `scrapedAt`

Missing values are omitted rather than fabricated. `thumbnail` is the first URL in `images` when media is enabled. Image output is restricted to public Alibaba CDN assets, with obvious logo, icon, avatar, placeholder, and loading assets removed. Operational counts and failure evidence are written to the `OUTPUT_SUMMARY` key-value record rather than mixed into product rows.

### Input

#### Keyword search

The simplest input creates a public Alibaba search URL:

```json
{
  "keyword": "wireless earbuds",
  "maxItems": 25,
  "maxPages": 3,
  "sortBy": "best_match"
}
```

#### Multiple category or search sources

Use explicit public URLs when you already have them. All sources share one global `maxItems` limit:

```json
{
  "startUrls": [
    "https://www.alibaba.com/trade/search?SearchText=wireless+earbuds&has4Tab=true&tab=all",
    "https://www.alibaba.com/trade/search?SearchText=bluetooth+speakers&has4Tab=true&tab=all"
  ],
  "maxItems": 50,
  "maxPages": 0,
  "deduplicate": true,
  "includeMedia": true,
  "includeDiagnostics": true
}
```

`maxPages: 0` means automatic pagination with a hard 100-page safety cap. Pages within one source are sequential; `maxConcurrency` controls how many source URLs can run at once.

#### Filters and sorting

```json
{
  "keyword": "led work light",
  "maxItems": 100,
  "maxPages": 5,
  "sortBy": "price_asc",
  "minPrice": 5,
  "maxPrice": 40,
  "minOrderQuantity": 10,
  "minSupplierRating": 4.5,
  "minReviewCount": 20,
  "minSoldCount": 100,
  "supplierCountry": "CN",
  "certifiedOnly": true,
  "requiredBadges": ["CE"],
  "excludeAds": true
}
```

Supported `sortBy` values are `best_match`, `price_asc`, `price_desc`, `rating_desc`, `reviews_desc`, `sold_desc`, and `supplier_years_desc`.

#### Developer and proxy options

```json
{
  "keyword": "wireless earbuds",
  "requestDelayMs": 1200,
  "maxConcurrency": 2,
  "maxRequestRetries": 3,
  "requestTimeoutSecs": 60,
  "requestHandlerTimeoutSecs": 300,
  "selectorTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

Proxy configuration changes only the network route. It does not bypass a challenge or access control. `includeDiagnostics: false` suppresses diagnostic rows in the dataset, but `diagnosticsFound` and the blocked or partial status remain in `OUTPUT_SUMMARY`.

#### Complete input reference

The supported canonical fields are:

- `keyword`: search phrase; defaults to `wireless earbuds` when no source URL is supplied.
- `startUrls`: up to 20 public HTTPS Alibaba URLs.
- `categoryUrl` and `categoryUrls`: convenience fields appended after `startUrls`.
- `maxItems`: global emitted product-row limit, 1–500; default 40.
- `maxPages`: pages per source, 0–100; default 3. Zero enables bounded automatic pagination.
- `sortBy`: `best_match`, `price_asc`, `price_desc`, `rating_desc`, `reviews_desc`, `sold_desc`, or `supplier_years_desc`.
- `includeMedia`, `deduplicate`, and `includeDiagnostics`: booleans, all defaulting to true.
- `excludeAds` and `certifiedOnly`: booleans, default false.
- `supplierCountry`, `supplierLocation`, `minPrice`, `maxPrice`, `minOrderQuantity`, `minSupplierRating`, `minReviewCount`, `minSoldCount`, and `requiredBadges`: optional filters.
- `requestDelayMs`: 0–15,000 ms; default 1,000.
- `maxConcurrency`: 1–5 source workers; default 2.
- `maxRequestRetries`: 0–10; default 3.
- `requestTimeoutSecs`: 15–300 seconds; default 60.
- `requestHandlerTimeoutSecs`: 60–900 seconds; default 300.
- `selectorTimeoutSecs`: 5–120 seconds; default 30.
- `proxyConfiguration`: optional `useApifyProxy`, `apifyProxyGroups`, and `apifyProxyCountry`.
- `fixtureFile`: an Actor-relative `.html` fixture for local QA only; Cloud runs reject it.

For compatibility, the runtime also accepts legacy aliases such as `keywords`, `searchQuery`, `query`, `queries`, `sourceUrl`, `sourceUrls`, `maxProducts`, `max_products`, `max_pages`, `pageDelayMs`, `navigationTimeoutSecs`, `timeoutSecs`, and `timeoutMs`. Canonical fields are recommended for new runs.

### Deterministic local QA

The bundled fixture is for local validation and cannot be used to create Cloud output:

```json
{
  "startUrls": ["https://www.alibaba.com/"],
  "maxItems": 1,
  "maxPages": 1,
  "includeMedia": true,
  "requestDelayMs": 0,
  "maxConcurrency": 1,
  "maxRequestRetries": 0,
  "requestTimeoutSecs": 15,
  "fixtureFile": "fixtures/sample.html"
}
```

For pagination fixtures, page 2 and later are loaded from sibling files named `sample-page-2.html`, `sample-page-3.html`, and so on.

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

### Run it in Apify Console

1. Open the Actor and select the **Input** tab.
2. Paste a canonical JSON input, such as the keyword example above, or enter the fields in the form.
3. Select **Save & Run** (or **Start**) and wait for the run to finish.
4. Open the **Dataset** tab to inspect product rows and any emitted diagnostics.
5. Open the **Key-value store** tab and inspect `OUTPUT_SUMMARY` for status, counts, pagination, filters, and failure evidence.
6. Use the Dataset export controls to download JSON, HTML, CSV, or Excel.

### Product output example

The following rich row mirrors the bundled public-shape fixture. Live values vary and are emitted only when observed in the public response:

```json
{
  "url": "https://www.alibaba.com/product-detail/rechargeable-led-work-light_1234567890123.html",
  "productId": "1234567890123",
  "productName": "Rechargeable LED Work Light",
  "title": "Rechargeable LED Work Light",
  "productUrl": "https://www.alibaba.com/product-detail/rechargeable-led-work-light_1234567890123.html",
  "price": 12.5,
  "priceCurrency": "USD",
  "priceFormatted": "12.50",
  "priceMin": 12.5,
  "priceMax": 12.5,
  "category": "Portable Lighting",
  "brand": "BrightForge",
  "availability": "https://schema.org/InStock",
  "supplierName": "BrightForge Industrial Co.",
  "supplierRatingValue": 4.8,
  "ratingsCountValue": 38,
  "images": ["https://s.alicdn.com/@sc04/kf/LED-work-light.jpg"],
  "thumbnail": "https://s.alicdn.com/@sc04/kf/LED-work-light.jpg",
  "imageCount": 1,
  "sourceUrl": "https://www.alibaba.com/",
  "pageNumber": 1,
  "position": 1,
  "scrapedAt": "2026-09-05T00:00:00.000Z"
}
```

### HTTP fallback output example

If browser extraction fails before producing records, the bounded HTTP fallback uses the same product-row contract. It does not add a fake transport field or claim fields that were not present:

```json
{
  "url": "https://www.alibaba.com/product-detail/example-public-item_123456789.html",
  "productId": "123456789",
  "productName": "Example public item",
  "productUrl": "https://www.alibaba.com/product-detail/example-public-item_123456789.html",
  "price": 4.25,
  "priceCurrency": "USD",
  "priceFormatted": "4.25",
  "sourceUrl": "https://www.alibaba.com/trade/search?SearchText=example",
  "pageNumber": 1,
  "position": 1,
  "scrapedAt": "2026-09-08T00:00:00.000Z"
}
```

### Diagnostic output example

Blocked or empty sources use a minimal four-field diagnostic shape:

```json
{
  "url": "https://www.alibaba.com/trade/search?SearchText=wireless+earbuds",
  "error": "Alibaba presented an access-control or challenge response.",
  "errorCode": "BLOCKED_SOURCE",
  "scrapedAt": "2026-09-08T00:00:00.000Z"
}
```

### Run summary example

`OUTPUT_SUMMARY` contains operational evidence without changing the dataset row contract:

```json
{
  "status": "SUCCEEDED",
  "keyword": "wireless earbuds",
  "sourceCount": 1,
  "maxItems": 25,
  "maxPages": 3,
  "pagesFetched": 2,
  "productCount": 2,
  "successfulCount": 2,
  "itemCount": 2,
  "diagnosticCount": 0,
  "diagnosticsFound": 0,
  "dataAvailable": true,
  "duplicateCount": 0,
  "filteredOutCount": 0,
  "usedProxy": false,
  "fallbackUsed": false,
  "finishedAt": "2026-09-08T00:00:00.000Z"
}
```

### Cost

The run uses Apify compute time and a Chromium browser. Enabling Apify Proxy can add proxy usage charges according to your Apify plan. Actual cost depends on memory, run duration, page count, retries, and proxy routing; review the run cost shown by Apify for the authoritative amount.

### Troubleshooting, API, and Issues

- `BLOCKED_SOURCE`, `HTTP_401`, `HTTP_403`, `HTTP_429`, `HTTP_451`, and `BROWSER_LAUNCH_FAILED` indicate a bounded access or runtime failure; inspect `failureEvidence` in `OUTPUT_SUMMARY`.
- `NO_PRODUCT_EVIDENCE` means the public response was readable but did not expose a supported Product, offer-list, SSR-card, or JSON-LD record.
- Try a smaller page limit, a respectful `requestDelayMs`, or a different public category URL. Proxy routing may change availability but is not an access-control bypass.
- Use the Actor **API** tab for the run endpoint and input schema. The Dataset API exposes product rows, while the key-value-store API exposes `OUTPUT_SUMMARY`.
- Report reproducible Actor problems through the Actor **Issues** tab with the run URL, sanitized input, and relevant `errorCode`; do not include credentials or private data.

### Privacy, legal, and responsible use

Use only public pages and information you are authorized to collect. Respect Alibaba's terms, robots guidance, applicable privacy and data-protection laws, rate limits, and any supplier or image rights. Do not use this Actor to collect private account data, bypass login or CAPTCHA controls, or create misleading commercial claims. You are responsible for validating supplier information and for how exported data is used.

### Run and validate locally

From this Actor directory:

```bash
npm install
npm test
apify validate-schema
apify run --purge --input-file .actor/input.json
npm run validate
```

`npm run validate` checks that product rows have canonical Alibaba product URLs, no debugging wrapper fields, safe product-scoped images, and the `thumbnail === images[0]` invariant. The summary distinguishes emitted diagnostics (`diagnosticCount`) from all observed failures (`diagnosticsFound`).

# Actor input Schema

## `keyword` (type: `string`):

Product phrase used to build the default Alibaba search URL. If startUrls is supplied, the URL's SearchText is used when this is omitted.

## `startUrls` (type: `array`):

One or more public HTTPS Alibaba category, search, catalog, or product-detail URLs. All sources share one maxItems limit.

## `categoryUrl` (type: `string`):

Convenience field for one Alibaba category/search URL; it is processed after startUrls.

## `categoryUrls` (type: `array`):

Additional Alibaba category/search URLs; they are processed after startUrls.

## `maxItems` (type: `integer`):

Global maximum number of product rows emitted across every source and page.

## `maxPages` (type: `integer`):

Pages per source. Use 0 for bounded automatic pagination up to 100 pages; every page remains subject to maxItems.

## `sortBy` (type: `string`):

Sort after extraction and filtering; best\_match preserves page and card order.

## `includeMedia` (type: `boolean`):

Keep product image URLs and thumbnail. Images are restricted to product-scoped public Alibaba CDN assets.

## `deduplicate` (type: `boolean`):

Remove repeated product IDs or canonical product URLs across pages and sources.

## `includeDiagnostics` (type: `boolean`):

Add minimal {url, error, errorCode, scrapedAt} rows when a source is blocked or has no public product evidence.

## `excludeAds` (type: `boolean`):

Exclude rows explicitly marked as advertisements by the public page.

## `certifiedOnly` (type: `boolean`):

Keep rows with an observed certification, verified, assessed, or equivalent public supplier signal.

## `supplierCountry` (type: `string`):

Optional two- or three-letter supplier country code, or any.

## `supplierLocation` (type: `string`):

Optional case-insensitive text match against observed supplier location or country.

## `minPrice` (type: `number`):

Keep rows whose observed price range reaches this value.

## `maxPrice` (type: `number`):

Keep rows whose observed price range reaches this value.

## `minOrderQuantity` (type: `number`):

Keep rows with an observed MOQ at or above this value.

## `minSupplierRating` (type: `number`):

Keep rows with an observed supplier/product rating at or above this value.

## `minReviewCount` (type: `integer`):

Keep rows with at least this observed rating/review count.

## `minSoldCount` (type: `integer`):

Keep rows with at least this observed sold count.

## `requiredBadges` (type: `array`):

Keep rows whose observed badge, certification, or selling-point text contains every value.

## `requestDelayMs` (type: `integer`):

Delay before page navigations and retries. Respectful pacing is recommended.

## `maxConcurrency` (type: `integer`):

Maximum number of source URLs processed concurrently; pages within one source remain sequential.

## `maxRequestRetries` (type: `integer`):

Retries for transient navigation failures. Challenge responses fail closed and are not bypassed.

## `requestTimeoutSecs` (type: `integer`):

Per-navigation timeout.

## `requestHandlerTimeoutSecs` (type: `integer`):

Maximum time budget documented for one source/page operation.

## `selectorTimeoutSecs` (type: `integer`):

Timeout used while waiting for the public page body before extracting HTML and embedded data.

## `proxyConfiguration` (type: `object`):

Optional standard Apify Proxy route for public requests. Proxying does not solve CAPTCHA or bypass access controls.

## `fixtureFile` (type: `string`):

Optional Actor-relative HTML fixture for deterministic QA. Page 2+ automatically uses sibling files named -page-2.html, -page-3.html, etc., when present.

## Actor input object example

```json
{
  "keyword": "wireless earbuds",
  "maxItems": 40,
  "maxPages": 3,
  "sortBy": "best_match",
  "includeMedia": true,
  "deduplicate": true,
  "includeDiagnostics": true,
  "excludeAds": false,
  "certifiedOnly": false,
  "requestDelayMs": 1000,
  "maxConcurrency": 2,
  "maxRequestRetries": 3,
  "requestTimeoutSecs": 60,
  "requestHandlerTimeoutSecs": 300,
  "selectorTimeoutSecs": 30
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing marketplace product facts and minimal blocked diagnostics, without Actor-debugging wrapper fields.

## `keyValueStore` (type: `string`):

Key-value store containing the OUTPUT\_SUMMARY run summary.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("w3crawler/alibaba-category-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("w3crawler/alibaba-category-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call w3crawler/alibaba-category-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,w3crawler/alibaba-category-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/t9ubJ5Q3yhv41o2hV/builds/XAxihFoUQ1v0OvUKt/openapi.json
