# Shopee Scraper (`mlg14/shopee-scraper`) Actor

Collect Shopee product records by keyword or numeric shop ID. When listing access is blocked, a Taiwan sitemap fallback returns product links, titles, and images without prices or seller metrics.

- **URL**: https://apify.com/mlg14/shopee-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.30 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Shopee Scraper

Collect public Shopee product listings by keyword, search URL, or numeric shop ID. The actor is intended to export product prices, images, sales counts, ratings, and seller indicators as structured records.

**Validation status:** the current public product endpoints returned HTTP 403 in the live golden check. That run produced zero records. This actor is not ready for publication or production use until a repeat golden check returns real products. Its output contract below describes the code path, not a verified production dataset.

### What data can you extract from Shopee?

Each successful record is one product. Product and shop identifiers together form the deduplication key. A value can be empty when a listing does not expose it. Prices are converted from the listing endpoint's hundred-thousandth currency units into the market's major currency unit. The current implementation does not request product detail pages, so it cannot promise descriptions, specifications, variant stock, or vouchers.

| Field | Description | Example or format |
| --- | --- | --- |
| `itemId` | Numeric product identifier | integer |
| `shopId` | Numeric shop identifier | integer |
| `url` | Direct product link in the selected market | product URL |
| `name` | Listing title | text |
| `image` | Main product image | image URL |
| `images` | Other image URLs exposed by the listing | array of URLs |
| `price` | Current listing price | local currency number |
| `priceMin` | Lowest listed variant price | local currency number |
| `priceMax` | Highest listed variant price | local currency number |
| `priceBeforeDiscount` | Former displayed price | local currency number |
| `discount` | Discount label returned by the listing | text |
| `currency` | Currency for the selected market | `TWD`, `SGD`, or another listed code |
| `soldMonthly` | Recent sales count supplied by the listing | integer |
| `historicalSold` | Historical sales count supplied by the listing | integer |
| `stock` | Listing stock if exposed | integer |
| `rating` | Average product rating | number |
| `ratingCount` | Number of ratings | integer |
| `likedCount` | Number of likes | integer |
| `shopName` | Public shop name when supplied | text |
| `shopLocation` | Public shop location | text |
| `isOfficialShop` | Official shop indicator | boolean |
| `isPreferredSeller` | Preferred seller indicator | boolean |
| `categoryId` | Listing category identifier | integer |
| `brand` | Brand label when supplied | text |
| `hasVideo` | Whether video metadata is present | boolean |
| `scrapedAt` | Record collection time | ISO 8601 timestamp |

The fields are deliberately flat. Product images are full URLs, while product and shop IDs remain separate so a downstream workflow can join records from multiple searches. An empty `priceBeforeDiscount` does not mean a discount exists but was missed; a listing may simply have no former price. An empty `shopName` means that name was absent from the listing response and should not be filled by guessing from a URL.

### How to scrape Shopee

1. Enter one or more product keywords in `searchTerms`, or provide a supported search URL or numeric shop ID.
2. Choose the market. The default is Taiwan. Use terms in the market's language where possible.
3. Set a maximum product count and, if useful, a sort order or price and rating filters.
4. Run the actor and inspect its status. A blocked listing request fails visibly. A successful run can then be exported from its dataset.

The actor combines results from every supplied keyword and shop source. It removes duplicate products using the pair of shop and item identifiers. `maxItems` is a run-wide cap, not a cap per keyword. If the first keyword fills the cap, later keywords are not requested. To compare the same number of records from each keyword, run each keyword separately.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `searchTerms` | string array | empty | Product keywords. |
| `startUrls` | URL list | empty | Search pages with a `keyword` query parameter or numeric `/shop/` pages. |
| `shopId` | integer | empty | Numeric shop catalog ID. |
| `region` | string | `TW` | `TW`, `SG`, `ID`, `MY`, `PH`, `TH`, `VN`, or `BR`. |
| `maxItems` | integer | 100 | Maximum unique products across the run, up to 5,000. |
| `sort` | string | `relevance` | Relevance, newest, ascending price, descending price, or best selling. |
| `priceMin` | number | empty | Minimum current price in local currency. |
| `priceMax` | number | empty | Maximum current price in local currency. |
| `minRating` | number | empty | Minimum average star rating. |
| `minSold` | integer | empty | Minimum historical sales count. |
| `discountedOnly` | boolean | false | Keep records with a discount label. |
| `officialShopOnly` | boolean | false | Keep official shop listings. |
| `preferredSellerOnly` | boolean | false | Keep preferred seller listings. |
| `withVideoOnly` | boolean | false | Keep listings containing video metadata. |
| `proxyConfiguration` | object | enabled | Connection settings. |

At least one of `searchTerms`, `startUrls`, and `shopId` is required. For a URL input, the hostname must match the selected region. A search URL needs a `keyword` query parameter. A shop URL needs a numeric ID in `/shop/<id>`. Product detail URLs and shop vanity names are outside the accepted URL format. Invalid sources fail before the first product request.

The golden input is:

```json
{
  "searchTerms": ["手機殼"],
  "region": "TW",
  "maxItems": 40,
  "sort": "relevance"
}
```

That check requires at least 30 results and requires IDs, URLs, names, currency, and strong price and image coverage. Its latest run returned zero products because the listing endpoint denied the request.

### Output example

There is no verified product record to show. The golden run returned zero records, so publishing an invented example here would misrepresent the available data. Once a golden run succeeds, replace this section with one actual dataset item and confirm the field fill rates against that item set.

### Use cases

- Track public listing prices for a defined set of product keywords over time.
- Compare the public assortment of several numeric shop catalogs.
- Monitor ratings and recent sales indicators in a market segment.
- Export public product links and images for catalog review.
- Find discounted listings within a price range.
- Compare visible official shop and preferred seller indicators.

These uses depend on obtaining records. The current HTTP 403 response prevents the actor from serving them. A scheduled workflow should treat a failed run as missing data, never as evidence that products disappeared from the market.

### How much does it cost to scrape Shopee?

The current build has not passed validation, so there is no reliable result-based cost or throughput measurement. The run attempted its first page and returned zero billable product records. Connection resources can still be consumed during failed attempts. Do not estimate a 1,000-product campaign from this failed run. Once access is restored, measure a successful 100-product run, then use the configured per-result price to estimate 1,000 and 5,000 products. For example, at a configured price of **P per 1,000 results**, 100 results would be **0.1 × P**, 1,000 would be **P**, and 5,000 would be **5 × P**, before any account-specific limits.

### Tips for best results

Use specific search terms in the language commonly used by sellers in the selected market. The search endpoint paginates in groups of 60; this actor requests at most 30 pages for each source. A broad keyword can therefore expose only the beginning of the site's ranked result set. Add narrower terms to reach different result groups, and use the deduplication behavior to combine overlapping searches safely.

Choose a sort order that matches the question being asked. Relevance is appropriate for discovering products, newest for recent listings, ascending price for inexpensive offers, and best selling for high-volume listings. Sort order changes the records exposed by the site; it does not guarantee that every matching product can be reached through pagination. Price, rating, sales, and seller filters are applied to returned listing data. A strict local filter may leave fewer than `maxItems` records even when the site has more matching products.

Keep market and URL host consistent. A Taiwan search URL belongs with `TW`; a Singapore search URL belongs with `SG`. The actor constructs product links and image URLs for the chosen market. Product prices use that market's local currency, so do not compare raw price numbers across markets without a separate currency conversion step.

### Limits

The principal limit is current access: public product search and shop listing endpoints returned HTTP 403 during research and the live golden run. The browser search page reached a traffic verification error. Changing sort parameters, normal request headers, proxy exit, and an app-style request did not produce product records. Public category and suggestion endpoints responded, but they do not contain the product listing fields needed here.

The listing endpoint is separate from product detail pages. Even after search access is restored, this implementation will not provide full descriptions, specifications, stock by variant, promotional voucher terms, or detailed shipping estimates. The table above is the implemented output contract. Some listing fields can be absent by product or market, and the actor represents them as empty values rather than inferred facts.

Pagination stops after 30 pages per keyword or shop. Results are deduplicated, so the final count can be smaller than the number of rows returned by the site. No login credentials are requested. If the site changes its response shape or access controls, the run can fail or produce fewer records and should be revalidated before use.

### Automation

A consumer can pass the input object to a scheduled run, read the dataset after success, and compare the latest records with a prior snapshot. It should check the run status and expected row count before replacing a previous snapshot. A zero-row or failed run must remain visible as an ingestion failure. Once the golden check passes, the dataset can feed spreadsheets, internal price reports, and alerting workflows.

### FAQ

#### Can I collect products by keyword and by shop in one run?

Yes. Supply search terms and a numeric shop ID together. Results are combined, and duplicate shop/item pairs are retained only once. The maximum count applies to the combined result set.

#### Can I paste a product URL?

No. The URL input accepts search pages with a keyword query and shop pages with numeric IDs. This actor collects listings rather than product detail pages.

#### Why is a field empty?

The listing response may not supply it for that product or market. Seller names, stock, brand, and former prices are examples of fields that can vary. Empty values should be treated as unknown, not as zero.

#### Does the actor need login credentials?

No credentials are accepted. It requests public listing data. At present, anonymous product requests are blocked even though some public metadata endpoints work.

#### How fast is a run?

No successful product run has been measured. The latest golden run failed on its first product page, so a speed estimate would be misleading.

#### Can results be exported to a spreadsheet?

After a successful run, the flat dataset can be exported in spreadsheet-compatible formats. No successful dataset exists yet for this build.

#### Is public product collection always permitted?

Only collect public data in a way that complies with the site's terms and applicable law. Avoid using the data to identify or profile people. This actor is scoped to product listings and public shop indicators.

### Integrations

Once the actor passes validation, its default dataset can be consumed through an API or scheduled export and connected to workflow systems. Check run success and record count before sending data onward. A failed collection should trigger an alert rather than overwrite a report with an empty table.

### Support

Open an issue describing the market, input, run status, and endpoint error. Include the count and field fill rates from a recent validation run when available.

# Actor input Schema

## `searchTerms` (type: `array`):

Product keywords to search. Results are combined and duplicate products are removed.

## `startUrls` (type: `array`):

Search URLs with a keyword parameter or numeric /shop/ URLs in the selected market.

## `shopId` (type: `integer`):

Numeric shop ID whose public product catalog will be collected.

## `region` (type: `string`):

Market to search. Prices use its local currency.

## `maxItems` (type: `integer`):

Maximum unique products across all searches and shops.

## `sort` (type: `string`):

Order requested from the listing endpoint.

## `priceMin` (type: `number`):

Include products priced at or above this amount in local currency.

## `priceMax` (type: `number`):

Include products priced at or below this amount in local currency.

## `minRating` (type: `number`):

Include products with at least this average star rating.

## `minSold` (type: `integer`):

Include products with at least this many recorded historical sales.

## `discountedOnly` (type: `boolean`):

Include only listings that show a discount.

## `officialShopOnly` (type: `boolean`):

Include only listings marked as official shops.

## `preferredSellerOnly` (type: `boolean`):

Include only listings with a preferred seller flag.

## `withVideoOnly` (type: `boolean`):

Include only listings that provide video information.

## `proxyConfiguration` (type: `object`):

Proxy configuration used for site requests.

## `sitemapFallback` (type: `boolean`):

For Taiwan searches without filters, return indexed product links, titles, and images when listing requests are blocked. Sitemap records have no price, rating, or sales data.

## Actor input object example

```json
{
  "searchTerms": [
    "手機殼"
  ],
  "region": "TW",
  "maxItems": 100,
  "sort": "relevance",
  "discountedOnly": false,
  "officialShopOnly": false,
  "preferredSellerOnly": false,
  "withVideoOnly": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "sitemapFallback": true
}
```

# Actor output Schema

## `products` (type: `string`):

Collected products in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "手機殼"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/shopee-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchTerms": ["手機殼"] }

# Run the Actor and wait for it to finish
run = client.actor("mlg14/shopee-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "手機殼"
  ]
}' |
apify call mlg14/shopee-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/shopee-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QcfSmwa6YbqpBmGSD/builds/RvFOAaSPgiVIvbo79/openapi.json
