# SHEIN Product Catalog – Fast Historical Data (`fetchfinch/shein-product-catalog`) Actor

\[ 💰 $0.70 / 1K]  Search 216,699 historical SHEIN catalog records and export useful product data at $0.70 per 1,000 records. Fast indexed lookups, keyword and category filters, product IDs, and archived USD/EUR prices. No proxy or API key required.

- **URL**: https://apify.com/fetchfinch/shein-product-catalog.md
- **Developed by:** [Fetch Finch](https://apify.com/fetchfinch) (community)
- **Categories:** E-commerce
- **Stats:** 8 total users, 4 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 product results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### Fast, affordable access to 216,000+ SHEIN catalog records

![Catalog](https://img.shields.io/badge/Catalog-216%2C699%20records-00897B)
![Pricing](https://img.shields.io/badge/Pricing-%240.70%20%2F%201%2C000%20records-5C6BC0)
![No API key required](https://img.shields.io/badge/No%20API%20key-required-43A047)
![Indexed search](https://img.shields.io/badge/Search-fast%20%26%20indexed-3178C6)
![Export](https://img.shields.io/badge/Export-JSON%20%7C%20CSV%20%7C%20Excel-F39C12)
![Examples](https://img.shields.io/badge/Examples-ready--to--run-00A98F)

Turn SHEIN catalog data into useful research datasets in a few clicks. Search **216,699 historical catalog records**, find products by ID or SKU, filter archived prices, and export exactly the records you need—without downloading and processing large source files yourself.

This Actor queries a bundled, indexed catalog. It does not need to open SHEIN, wait for a scraping provider, or solve website verification challenges. No proxy setup, external API key, or additional scraping-provider subscription is required.

**Data type:** historical snapshots, not a live product feed. Use it for research, analysis, and prototyping; current prices and stock require a live source.

### Why choose this catalog?

- **Fast access:** indexed keyword and product lookups run locally, with no upstream scraping queue.
- **Low-cost exports:** $0.70 per 1,000 returned records, plus the standard small start charge.
- **Broad research coverage:** explore clothing, accessories, beauty, footwear, electronics, and other product groups across four source snapshots.
- **Useful product information:** titles, IDs, SKUs, attributes, image links, size labels, and archived prices where the source supplies them.
- **Flexible selection:** combine search terms with source, market, category, currency, and price filters.
- **Ready for your workflow:** use Apify datasets, JSON/CSV/Excel exports, APIs, webhooks, spreadsheets, or no-code integrations.
- **Traceable data:** every record identifies its source snapshot, catalog version, and collection date where known.

### What can you build with it?

**Product research and assortment analysis:** examine archived product names, categories, attributes, and price points.

**Search and recommendation prototypes:** develop catalog browsing, keyword discovery, and classification workflows using structured product records.

**Ecommerce integration tests:** populate sample catalogs and test importers, exports, spreadsheets, or downstream automations.

**Historical datasets:** collect focused subsets for analysis without repeatedly cleaning and filtering the original CSV files.

### Pricing

The current price is **$0.70 per 1,000 dataset records** (`$0.0007` each), plus **$0.00005 per Actor start** at the default memory allocation. For example, 100 returned records correspond to $0.07 in record charges, plus the start charge. Consult the Pricing tab for the active billing terms.

Only returned catalog records are written to the dataset. Empty results and validation errors do not generate paid placeholder records; the start charge still applies. You can set a maximum spending limit, and the Actor respects the remaining dataset-item allowance.

### Ready-to-run examples

Copy one of the configurations below into the Actor's input to try these workflows:

- **Bluetooth headphones:** explore a small electronics research sample from the all-category snapshot.
- **Product lookup:** retrieve a known product ID and its archived metadata.
- **US dress research:** export matching product links, price points, and available attributes.
- **German beauty under €10:** filter the January 2023 snapshot by its archived EUR price.

### How to use it

1. Enter a search phrase, product IDs, or filters. Empty input returns a 20-record sample.
2. Choose `results_wanted`, up to 10,000 records per run.
3. Start the Actor and inspect the dataset preview.
4. Download your export or connect the dataset to your workflow.

#### Search the catalog

```json
{
  "query": "dress",
  "market": "us",
  "productUrlFilter": "with-url",
  "results_wanted": 100
}
```

#### Filter archived beauty prices

```json
{
  "sourceDataset": "de-beauty",
  "currency": "EUR",
  "maxPrice": 10,
  "results_wanted": 100
}
```

#### Explore Bluetooth headphone records

```json
{
  "query": "bluetooth headphones",
  "sourceDataset": "all-categories",
  "results_wanted": 20
}
```

This research snapshot provides archived titles and prices without product IDs or URLs.

#### Look up a known product

```json
{
  "productIds": ["12979206"],
  "market": "us",
  "results_wanted": 1
}
```

### Input parameters

| Parameter | Default | What it does |
| --- | --- | --- |
| `query` | — | Searches historical titles and descriptions by word prefix. |
| `searchTerms` | `[]` | Adds up to 20 search phrases. Phrases use OR; words within each phrase use AND. |
| `productIds` | `[]` | Looks up up to 100 numeric SHEIN IDs, supplied as strings. |
| `skus` | `[]` | Looks up up to 100 exact SKU values, without the `SKU:` prefix. |
| `category` | — | Filters an exact source category label, ignoring ASCII case. Available labels are in `CATALOG_MANIFEST`. |
| `categoryId` | — | Filters the historical SHEIN category ID where available. |
| `market` | `all` | Selects `all`, `us`, `de`, or `unknown` source coverage. |
| `sourceDataset` | `all` | Selects one of the four source snapshots below, or all of them. |
| `currency` | `all` | Selects `USD`, `EUR`, or all currencies. Required for price filters. |
| `minPrice`, `maxPrice` | — | Filters archived prices in the selected currency. |
| `productUrlFilter` | `all` | `all`: all records; `with-url`: linked products; `without-url`: unlinked research records. |
| `results_wanted` | `20` | Limits returned records, subject to your spending limit. |
| `offset` | `0` | Skips matching records for the next export page. |

Filters combine with AND. Search is word-prefix matching: `dress` can match `dresses`. Related words, spelling variants, and languages are not automatically expanded. Results follow stable catalog order rather than relevance order. USD and EUR are not converted.

JSON integrations can also use the boolean `hasProductUrl` alias: `true` selects linked records and `false` selects unlinked records. The form uses the explicit selector above.

### Catalog coverage

| Snapshot | Records | Market / currency | Collection date | Available information |
| --- | ---: | --- | --- | --- |
| `us-catalog` | 103,202 | US / USD | Unknown | Broad catalog with product IDs, titles, attributes, size labels, and image links where supplied. |
| `de-beauty` | 30,255 | Germany / EUR | January 29, 2023 | Beauty titles, IDs, SKUs, category labels, and archived prices. |
| `us-footwear` | 1,151 | US / USD | Unknown | Additional footwear records, product links, and available image/price information. |
| `all-categories` | 82,091 | Unknown / USD | Unknown | Broad product-title, category, and price research records. |

Records with known product IDs are deduplicated within each market. The all-category snapshot lacks product IDs, SKUs, and product URLs, and may overlap other snapshots; the total is a catalog-record count, not a guarantee of unique products across all sources. German beauty records do not include images.

### Output data

Each dataset item contains a structured historical record:

| Fields | Information |
| --- | --- |
| `recordId`, `productId`, `sku`, `title`, `url` | Catalog identity and archived product details. |
| `description`, `brand`, `sizes`, `category`, `categoryId` | Attributes and categorization where supplied. |
| `image`, `images` | Archived image links. |
| `historicalPrice`, `historicalRetailPrice`, `priceText`, `priceIsFrom`, `currency` | Archived numeric prices and original price labels. |
| `sourceDataset`, `market`, `collectionDate` | Source provenance; unknown dates are `null`. |
| `isLive`, `snapshotVersion`, `retrievedAt` | Historical-data flag, catalog version, and export timestamp. |

For example, the product lookup task returns fields such as:

```json
{
  "recordType": "historicalProduct",
  "recordId": "us:12979206",
  "productId": "12979206",
  "title": "Flower Decor Face Covering Chain",
  "url": "https://us.shein.com/Flower-Decor-Face-Covering-Chain-p-12979206-cat-3028.html",
  "sku": "sc2302055742295937",
  "historicalPrice": 2.5,
  "currency": "USD",
  "sizes": "one-size",
  "sourceDataset": "us-catalog",
  "collectionDate": null,
  "isLive": false
}
```

Additional metadata varies by source. Missing fields remain `null` or empty rather than being inferred. Apify provides JSON, CSV, Excel, XML, and other export formats automatically.

### Speed, pagination, and integrations

Catalog queries use a prebuilt search index instead of parsing the source CSV files on every run. Small queries are designed to finish quickly; container startup and dataset storage still affect the full run time. Larger exports involve more storage and data transfer.

Use `RUN_DIAGNOSTICS.nextOffset` as your next `offset`, keeping the same filters and `snapshotVersion`. A `null` next offset means the export is complete. Diagnostics also include match counts, saved records, and query duration.

Use the Apify API, webhooks, Make, Zapier, Google Sheets, Airtable, or your own code to consume results. Schedules are available, but repeating the same query against the same snapshot returns the same historical information.

### Data freshness and migration

This catalog is intended for informational and historical use. Prices, size labels, product links, and image links may have changed since collection; they are not current stock or purchasing information. `retrievedAt` is the export time, not the original collection date. Dates are reported only when the source provides a verified date; archive timestamps are not treated as collection dates.

This Actor has transitioned from live scraping to catalog search. The input examples above use catalog settings. For older custom integrations, replace URL, locale, listing-pagination, proxy, enrichment, and HTML settings with the input fields above. The legacy numerical `maxProducts` alias remains available for `results_wanted`.

### Support and responsible use

Use the Issues tab for query problems, feature requests, or requests for additional licensed snapshots. Use the data responsibly and comply with applicable rights and restrictions. This Actor is not affiliated with or endorsed by SHEIN.

# Actor input Schema

## `query` (type: `string`):

Words are matched as prefixes in historical titles and descriptions. All words in one query must match.

## `searchTerms` (type: `array`):

Up to 20 phrases, combined with OR. All words within each phrase must match.

## `productIds` (type: `array`):

Up to 100 numeric SHEIN IDs as strings. Only sources with known IDs can match.

## `skus` (type: `array`):

Up to 100 exact SKU values, without the SKU: prefix.

## `category` (type: `string`):

Exact category label, ignoring ASCII case. Available labels are listed in CATALOG\_MANIFEST. Not all sources contain category names.

## `categoryId` (type: `string`):

Exact historical SHEIN category identifier where available.

## `market` (type: `string`):

The all-category file has no verified market and is labeled unknown.

## `sourceDataset` (type: `string`):

Collection dates are unknown except for German beauty (2023-01-29).

## `currency` (type: `string`):

Historical source currency. Required for price filters. No currency conversion.

## `minPrice` (type: `number`):

Filter archived prices, not current prices. Select a currency.

## `maxPrice` (type: `number`):

Filter archived prices, not current prices. Select a currency.

## `productUrlFilter` (type: `string`):

Include all records, only records with archived product links, or only unlinked research records.

## `results_wanted` (type: `integer`):

Maximum historical records to save, subject to your spending limit.

## `offset` (type: `integer`):

Skip this many matching records. Use RUN\_DIAGNOSTICS.nextOffset with identical filters and snapshot version for the next export.

## Actor input object example

```json
{
  "searchTerms": [],
  "productIds": [],
  "skus": [],
  "market": "all",
  "sourceDataset": "all",
  "currency": "all",
  "productUrlFilter": "all",
  "results_wanted": 20,
  "offset": 0
}
```

# Actor output Schema

## `products` (type: `string`):

Search results with historical prices, source snapshot, collection date where known, and isLive=false.

## `diagnostics` (type: `string`):

Match count, saved count, pagination offset, runtime, and validation errors.

## `manifest` (type: `string`):

Snapshot version, source checksums, collection-date limitations, counts, and category labels.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("fetchfinch/shein-product-catalog").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("fetchfinch/shein-product-catalog").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call fetchfinch/shein-product-catalog --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fetchfinch/shein-product-catalog"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/h10NQBck1kjH7pIGg/builds/HSrWqMpadDtDW9mtI/openapi.json
