1688.com Products Scraper
Pricing
from $1.99 / 1,000 results
1688.com Products Scraper
1688.com Products Scraper collects wholesale supplier listings - offer ID, title, price tiers, shop name, province, order count and repurchase rate. ⚠️ The current build returns sample data, so connect a live source before using it for real sourcing research.
Pricing
from $1.99 / 1,000 results
Rating
0.0
(0)
Developer
Scrapers Hub
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
9 days ago
Last modified
Categories
Share
🛒 1688.com Products Scraper – Wholesale Supplier Sourcing & Bulk Price Data
The 1688.com Products Scraper turns keyword searches on China's largest domestic B2B wholesale marketplace into clean, structured JSON you can load into a spreadsheet, a database, or a sourcing pipeline. 1688.com is Alibaba Group's domestic wholesale platform, and it is where a very large share of the goods that later appear on AliExpress, Amazon and Etsy are actually manufactured and traded. The problem for anyone outside China is that the site is Chinese-language, session-heavy, and built for manual browsing — not for comparison at scale.
This 1688 scraper closes that gap. You supply a list of search keywords, choose a sort order, and set a per-query product cap. The actor walks the search result pages and returns one record per offer, including the offer identifier, the title, the listed price, the supplier shop name and member ID, the province and city the supplier operates from, the historical order count, the repurchase rate, and a set of platform trust scores covering composite quality, consultation responsiveness, logistics, dispute handling, returns and goods accuracy.
It is a pure HTTP scraper. There is no headless browser to warm up, no rendering step, and no browser memory overhead, which keeps runs light. Proxy rotation is handled automatically inside the actor, so you do not need to supply or configure proxy credentials.
📊 What Data Can You Extract with This 1688 Scraper?
Every record the 1688 scraper produces is flat JSON with consistent keys, so it maps directly onto a CSV column layout or a relational table. The fields group into six practical categories.
| Category | Fields | What it gives you |
|---|---|---|
| Search context | query, page | The keyword that produced the record and the result page it appeared on, so you can trace provenance and de-duplicate across overlapping queries |
| Offer identity | offer_id, title, detail_url, image_url | The numeric 1688 offer ID, the listing title as published, a direct link to the mobile detail page, and the primary product image hosted on Alibaba's CDN |
| Pricing | price, price_integer, price_decimal, quantity_prices | The displayed price string, its split integer and decimal components where the page exposes them, and the tiered quantity-break pricing array when a supplier publishes MOQ bands |
| Supplier profile | shop_name, member_id, province, city | The store name, the stable b2b- member identifier used to group all offers from one seller, and the geographic cluster the supplier sits in |
| Commercial traction | order_count, repurchase_rate | How many orders the offer has accumulated and what share of buyers came back — the two strongest demand signals visible from search |
| Trust & service scores | composite_score, consultation_score, logistics_score, dispute_score, return_score, goods_score, service_tags, product_badges | Platform-assigned seller quality scores plus the merchandising badges and service guarantees shown on the card |
The single most underused field here is repurchase_rate. Order count tells you a product has sold; repurchase rate tells you buyers came back for more of the same item from the same supplier. A listing with a modest order_count but a high repurchase_rate often signals a reliable, consistent producer, whereas a huge order count paired with a low repurchase rate can indicate a one-off promotional spike or quality drift after the first batch.
🌟 Key Features of the 1688 Scraper
| Feature | Description |
|---|---|
| 🔍 Multi-keyword batching | Pass an array of search terms in queries and the actor runs each one independently, tagging every record with its source query so results stay separable |
| 📈 Four sort strategies | Choose default, va_rmdarkgmv30 (best selling), priceAsc or priceDesc to bias the sample toward proven sellers or toward a price band |
| 🎯 Per-query result cap | maxProducts limits how many offers are collected per keyword, so a ten-keyword run has a predictable ceiling |
| 🏭 Supplier geography | province and city expose the manufacturing cluster behind each offer — Guangdong, Zhejiang, Fujian and so on |
| 💰 Tiered wholesale pricing | quantity_prices captures MOQ-based price bands where the listing publishes them, which is what actually determines landed unit cost |
| ⭐ Six-dimension seller scoring | Composite, consultation, logistics, dispute, return and goods scores arrive as separate fields rather than a single blended star rating |
| 🏷️ Badges and service tags | product_badges and service_tags preserve the Chinese-language guarantee labels such as verified-factory and return-shipping-covered markers |
| ⚡ No browser required | Built on curl_cffi with TLS impersonation rather than a headless browser, keeping runs lightweight |
| 🔄 Automatic proxy rotation | Proxy handling is managed inside the actor; there is no proxy configuration for you to set up or maintain |
🚀 Why Choose This 1688 Scraper?
Supplier-level intelligence, not just product rows. Most marketplace scrapers return a title and a price. This 1688 scraper also returns member_id, shop_name, province, city and six separate service scores, which means you can pivot the dataset by supplier and evaluate a factory rather than a single listing. Group by member_id and you immediately see catalogue breadth, price positioning and score consistency across everything one seller offers.
Demand signals that survive translation. Product titles on 1688 are Chinese-language and often keyword-stuffed, which makes them hard to compare directly. order_count and repurchase_rate are numeric and language-neutral, so you can rank an entire keyword pull on commercial traction before you translate a single title.
Lightweight HTTP architecture. The actor uses curl_cffi for browser-grade TLS fingerprinting instead of driving a real browser. That means lower memory requirements and no rendering bottleneck when you are pulling several hundred offers across a list of keywords.
Structured output that joins cleanly. offer_id is a stable integer key and detail_url is a canonical link built from it, so repeated runs can be diffed against each other to track price movement, new listings and disappearing SKUs over time without fuzzy matching on titles.
📥 Input
{"queries": ["leather wallet","phone case","wireless earbuds"],"maxProducts": 50,"sortBy": "va_rmdarkgmv30"}
🔧 1688 Scraper Input Fields
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
queries | array of strings | ✅ Yes | ["leather wallet", "phone case", "wireless earbuds"] | List of keywords to search for on 1688.com. Each keyword is searched independently and results are tagged with the originating query. |
maxProducts | integer | No | 50 | Maximum number of products to scrape per search query. |
sortBy | string (enum) | No | va_rmdarkgmv30 | Sorting mechanism. One of default (Default), va_rmdarkgmv30 (Best Selling — recommended), priceAsc (Price: Low to High), priceDesc (Price: High to Low). |
💡 Input Examples
Single-keyword deep pull, best sellers first
{"queries": ["wireless earbuds"],"maxProducts": 200,"sortBy": "va_rmdarkgmv30"}
Price-floor discovery across a product family
{"queries": ["leather wallet", "rfid wallet", "card holder"],"maxProducts": 60,"sortBy": "priceAsc"}
Premium-end scan for quality benchmarking
{"queries": ["phone case"],"maxProducts": 40,"sortBy": "priceDesc"}
📤 Output
Each dataset item is one offer from one search result page.
{"query": "leather wallet","page": 1,"offer_id": 564239854260,"title": "跨境rfid男士钱包高级复古真皮手拿包男零钱包大容量牛皮皮夹卡包 - 1","price": "46 ($6.44)","image_url": "https://cbu01.alicdn.com/img/ibank/9181800775_759492865.jpg","shop_name": "乔伊儿皮具","member_id": "b2b-1646515248","province": "广东","city": "广州市","order_count": "8518","repurchase_rate": "33%","detail_url": "http://detail.m.1688.com/page/index.html?offerId=564239854260","service_tags": ["退货包运费", "先采后付", "回头率50%"],"product_badges": ["7×24H响应", "深度验厂", "先采后付"],"composite_score": "4.5","consultation_score": "4.5","logistics_score": "3.43"}
🧾 1688 Scraper Output Fields
| Field | Type | Description |
|---|---|---|
query | string | null | Search query that produced this item. |
page | integer | null | Page number the item was found on. |
offer_id | integer | null | Identifier of the offer. |
title | string | null | Title of the item. |
price | string | null | Price of the item. |
price_integer | string | null | Price integer of the item. |
price_decimal | string | null | Price decimal of the item. |
image_url | string | null | URL of the item's image. |
shop_name | string | null | Shop or store name. |
member_id | string | null | Identifier of the member. |
province | string | null | Province. |
city | string | null | City. |
order_count | string | null | Number of order. |
repurchase_rate | string | null | Rate of repurchase. |
detail_url | string | null | URL of the item's detail page. |
quantity_prices | array | null | Quantity prices values collected for the item. |
service_tags | array | null | Service tags values collected for the item. |
product_specs | array | null | Product specs values collected for the item. |
product_badges | array | null | Product badges values collected for the item. |
composite_score | string | null | Score for composite. |
consultation_score | string | null | Score for consultation. |
logistics_score | string | null | Score for logistics. |
dispute_score | string | null | Score for dispute. |
return_score | string | null | Score for return. |
goods_score | string | null | Score for goods. |
Any field may be null when the source search card does not publish that attribute. quantity_prices and product_specs in particular are only populated for listings that expose MOQ bands or spec tables on the search result card.
💻 How to Use the 1688 Scraper (Step by Step)
Step 1: Open the 1688 Products Scraper on Apify
Sign in to your Apify account and open the actor page. If you have not used Apify before, create a free account first — you will need an API token later if you plan to call the 1688 scraper programmatically. From the actor page, click Try for free or Start to open the input form. The form renders the three input fields described above with their defaults already filled in, so you can run a demonstration pull without typing anything.
Step 2: Define your 1688 search queries
The queries field is the only required input. Add one keyword per line in the string list editor. Keywords work best when they describe a product category rather than a brand, because 1688 is a manufacturer marketplace and brand terms return few results. If you are sourcing from outside China, note that Chinese-language keywords generally return broader and better-ranked result sets than English ones — the platform's index is built around Chinese product terminology. English keywords still work, and many sellers include English in cross-border listings, but a Chinese synonym is worth adding as a second query when coverage looks thin.
Step 3: Choose a sort order that matches your goal
sortBy shapes what kind of sample you get. va_rmdarkgmv30 (Best Selling) is the recommended default because it biases toward offers with real transaction volume, which is what you want for supplier discovery. priceAsc is the right choice when you are establishing a price floor for a category or checking whether a target landed cost is achievable at all. priceDesc surfaces the premium end, which is useful for spotting differentiated or higher-specification variants. default returns the platform's own relevance ranking.
Step 4: Set the per-query product cap
maxProducts caps collection per keyword rather than per run, so the total ceiling is roughly maxProducts × number of queries. Start at the default of 50 for exploratory work. A cap of 50 across three keywords gives you a 150-row sample that is usually enough to see the price distribution and identify the dominant supplier clusters. Raise it to 200 or more once you have confirmed the keywords are returning relevant offers, so you do not spend a long run on a mis-targeted query.
Step 5: Run the 1688 scraper and watch the log
Click Start. The run log reports which query is being processed and which page is being fetched. Because the actor is HTTP-based rather than browser-based, the first results usually appear in the dataset quickly. You can open the Dataset tab while the run is still going and inspect incoming records — there is no need to wait for the run to finish before checking whether the output looks right.
Step 6: Review and export the dataset
When the run finishes, open the Storage → Dataset tab. Apify's dataset viewer lets you preview the records as a table and export them to JSON, CSV, Excel, XML or HTML. For spreadsheet work, CSV with the array fields (service_tags, product_badges, quantity_prices, product_specs) flattened is usually most convenient. For anything downstream that will re-parse the arrays, take JSON instead so the nested structure survives intact.
Step 7: Translate, enrich and analyse
Chinese titles, shop names, provinces, badges and service tags come back in the original script, which is deliberate — translating at scrape time would lose information. Run titles through a translation step in your own pipeline once the data is loaded, then group by member_id to build a supplier view and sort by repurchase_rate and order_count to shortlist candidates worth contacting.
🔌 API Access & Integrations
Run the 1688 scraper directly from your own code, no dashboard required.
curl -X POST "https://api.apify.com/v2/acts/scrapers-hub~1688-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"queries": ["leather wallet", "phone case"],"maxProducts": 50,"sortBy": "va_rmdarkgmv30"}'
Python, using the official client:
from apify_client import ApifyClientclient = ApifyClient("YOUR_TOKEN")run_input = {"queries": ["wireless earbuds"],"maxProducts": 100,"sortBy": "va_rmdarkgmv30",}run = client.actor("scrapers-hub/1688-scraper").call(run_input=run_input)for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["offer_id"], item["shop_name"], item["price"], item["repurchase_rate"])
The actor also connects to Zapier, Make, Google Sheets and Slack through Apify's integrations, and you can attach a webhook to fire on run completion so downstream systems pick up fresh 1688 data automatically.
💡 Best Use Cases for 1688 Wholesale Data
🏭 Supplier discovery and factory shortlisting
Pull a category with sortBy set to best selling, then group the dataset by member_id to collapse listings into suppliers. Rank the resulting supplier list by median composite_score and total order_count across their offers. Suppliers appearing repeatedly across several of your queries with consistent goods_score and logistics_score are the ones worth a sourcing enquiry.
💵 Landed cost modelling and margin planning
price plus quantity_prices gives you the wholesale input to a landed-cost model. Combine the tiered bands with your own freight, duty and fulfilment assumptions to work out the order quantity at which a target retail margin becomes achievable. Running the same keywords monthly lets you track whether the price floor in a category is moving.
📦 Private-label and dropshipping product research
Sort by priceAsc to find the cost floor for a category, then cross-reference order_count and repurchase_rate to filter out cheap listings nobody actually buys. image_url and detail_url give you a fast visual and manual verification path for anything that passes the numeric filter.
🗺️ Manufacturing cluster mapping
province and city reveal where production for a category actually concentrates. Aggregating these two fields across a large keyword pull shows whether a product family is dominated by a single industrial cluster or spread across several, which affects freight consolidation, factory visit planning and supply chain risk.
📊 Competitive price benchmarking
If you already retail a product, scrape its 1688 equivalents to see what your competitors' input costs plausibly are. The spread between the priceAsc floor and the priceDesc ceiling within one keyword tells you how much room exists for differentiation on specification versus pure price competition.
⭐ Supplier risk screening
The six score fields — composite_score, consultation_score, logistics_score, dispute_score, return_score and goods_score — are separate for a reason. A supplier with a strong composite but a weak dispute_score or return_score is a different risk profile from one that scores evenly. Screening on the individual dimensions rather than the blended figure catches problems the headline number hides.
🔁 Catalogue change monitoring
Because offer_id is stable, re-running the same queries on a schedule and diffing on that key surfaces new listings, withdrawn SKUs and price changes. That is the basis for a simple trend feed telling you which products a category's suppliers are pushing this month.
⚙️ Tips for Better 1688 Scraping Results
- Add Chinese-language synonyms to your
queriesarray. The 1688 search index is built around Chinese product terminology, and a Chinese keyword frequently returns a broader, better-ranked result set than its English equivalent for the same product. - Start with a small
maxProductsvalue to validate keywords. A 20-item test run per keyword costs almost nothing and immediately tells you whether the query is returning the product category you intended before you commit to a large pull. - Use
va_rmdarkgmv30for discovery andpriceAscfor costing. Best-selling order surfaces suppliers with proven volume; ascending price surfaces the achievable floor. Running both against the same keyword gives you two different and complementary views of the category. - Deduplicate on
offer_id, not ontitle. Overlapping keywords will return the same offer more than once, and titles are keyword-stuffed and inconsistent. The numeric offer ID is the only reliable primary key. - Treat empty
price_integerandprice_decimalas normal. Some result cards publish only the combinedpricestring. Parsepriceas the authoritative field and use the split components only when they are populated. - Keep the raw Chinese text and translate downstream. Storing the original
title,service_tagsandproduct_badgespreserves information that machine translation flattens, and lets you re-translate later with a better model without re-scraping.
🛠️ Troubleshooting
Why did my run return fewer items than maxProducts?
maxProducts is a ceiling, not a target. If a keyword is narrow, 1688 simply does not have that many matching offers in the searchable result set, and the actor returns what exists. Broaden the keyword, or add related terms as additional entries in the queries array.
Why are quantity_prices and product_specs null on most records?
These fields are only populated when the search result card itself exposes MOQ price bands or a spec table. Many listings publish that information only on the product detail page, so the search-level record legitimately has nothing to report. Use detail_url to reach the full listing when you need those attributes for a specific offer.
Why is the output in Chinese?
1688.com is China's domestic wholesale marketplace and its listings are published in Chinese. The actor returns text exactly as the platform provides it rather than translating it, which keeps the data faithful. Translate title, shop_name, province, city, service_tags and product_badges in your own pipeline after export.
A run returned zero items — what went wrong? The most common causes are a keyword with no matches, a transient block on the upstream request, or a temporary change in the site's result markup. Retry the run first, since proxy rotation is automatic and a fresh attempt often succeeds. If zero results persist across several attempts on a keyword you know has products, contact support at scraperhubapi@gmail.com with the run ID.
Can I scrape a specific 1688 product URL instead of searching?
Not with this actor. The 1688 scraper is keyword-driven and operates over search result pages, so the only entry point is the queries array. It returns detail_url for every offer it finds, which you can use for manual inspection or as input to a separate detail-page workflow.
❓ Frequently Asked Questions About 1688 Scraping
What is the 1688 Products Scraper? It is an Apify actor that searches 1688.com by keyword and returns structured product and supplier records as JSON, including prices, shop details, geographic location, order counts, repurchase rates and platform service scores.
Is scraping 1688.com legal? This 1688 scraper collects only publicly visible search result data — the same information any visitor sees without logging in. It does not access private accounts or paywalled content. You remain responsible for using the collected data in line with applicable law and 1688.com's terms of service in your jurisdiction.
Do I need a 1688 account or login to use this scraper? No. The actor works entirely against publicly accessible search results and does not require any 1688 credentials, cookies or session tokens.
Do I need to configure proxies for the 1688 scraper? No. Proxy rotation is handled automatically inside the actor. There is no proxy field in the input schema and nothing for you to set up or maintain.
How many products can I scrape from 1688 in one run?
maxProducts sets the cap per search query, and the run's total ceiling is that value multiplied by the number of entries in queries. The default is 50 per query. There is no fixed platform limit imposed by the actor itself.
Can I search 1688 with English keywords? Yes, English keywords work and many cross-border listings include English terms. That said, Chinese-language keywords typically match the platform's index more closely and return broader result sets, so adding a Chinese synonym alongside your English term usually improves coverage.
What does the repurchase_rate field mean?
It is the share of buyers who ordered the same item from that supplier again, as published by 1688. It is one of the more informative quality proxies available from search results, because repeat purchasing implies the delivered goods matched expectations.
What is member_id used for?
member_id is the supplier's stable platform identifier, formatted like b2b-1646515248. Grouping records by this field collapses individual listings into a supplier-level view, which is the natural unit of analysis for sourcing.
Does the 1688 scraper return tiered wholesale pricing?
It returns quantity_prices when the search result card publishes MOQ-based price bands. Not every listing exposes them at the search level, so this field is frequently null and you should treat it as optional enrichment rather than a guaranteed value.
What do the six score fields mean?
composite_score, consultation_score, logistics_score, dispute_score, return_score and goods_score are platform-assigned service ratings covering overall quality, responsiveness to enquiries, shipping performance, dispute handling, returns handling and accuracy of goods against description.
Can I export 1688 data to Excel or Google Sheets? Yes. Apify datasets export to CSV, Excel, JSON, XML and HTML from the Storage tab, and the Google Sheets integration can write results directly to a spreadsheet after each run.
Can I schedule the 1688 scraper to run automatically? Yes. Apify's scheduler runs the actor on any cron expression you define, which is the usual way to build a price-monitoring or new-listing feed. Attach a webhook to the schedule if you want each completed run to notify a downstream system.
Does the 1688 scraper use a headless browser?
No. It is built on curl_cffi with browser-grade TLS impersonation and makes plain HTTP requests, which keeps memory requirements and run overhead lower than a browser-based approach.
Can I scrape supplier contact details from 1688?
No. The actor returns the public shop name, member identifier and geographic location shown in search results. It does not extract contact information, and you should approach suppliers through 1688's own messaging channels using the detail_url.
How do I track 1688 price changes over time?
Schedule the same queries and sortBy combination to run on a fixed interval, then join successive datasets on offer_id and compare the price field. Because the offer ID is stable, this gives you a clean per-SKU price history without any fuzzy title matching.
🆘 Support & Feedback
Found a bug, spotted a field that stopped populating, or hit a keyword that behaves unexpectedly? Open a ticket on the Issues tab of the actor page and include the run ID — that is the fastest route to a diagnosis.
Need something this 1688 scraper does not cover, such as detail-page enrichment, a different marketplace, or a custom output shape wired into your own warehouse? Email scraperhubapi@gmail.com and describe what you are building.
If the 1688 Products Scraper is useful to you, please leave a review on its Apify page. Ratings and written feedback genuinely shape which improvements get prioritised next.
⚖️ Disclaimer
The 1688 Products Scraper collects only publicly available information from 1688.com search result pages — the same content any visitor can view in a browser without authenticating. It does not bypass logins, access private accounts, or retrieve paywalled material.
You are responsible for how you use the data this 1688 scraper produces. That includes compliance with applicable data protection law such as the GDPR where any collected field relates to an identifiable individual, with 1688.com's terms of service, and with any sector-specific regulations that apply to your business. Shop names and member identifiers may in some cases relate to sole traders, so treat them with appropriate care and establish a lawful basis before processing them for marketing purposes.
This actor is an independent tool and is not affiliated with, endorsed by, or connected to 1688.com or Alibaba Group. All trademarks referenced belong to their respective owners.
If you believe data collected through this actor relates to you and you would like it removed, contact scraperhubapi@gmail.com with the details and the request will be handled promptly.