# Sephora Products Scraper (`scrapyx/sephora-products-scraper`) Actor

Every product in a Sephora US category: brand, name, price or price range in USD, sale price, rating and review count, bestseller / new / Sephora-exclusive flags, variants, image and product URL. HTTP only, no login, no proxy.

- **URL**: https://apify.com/scrapyx/sephora-products-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** E-commerce, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.56 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Sephora Products Scraper (US)

Every product in a **Sephora US category**: brand, product name, price or
price range in USD, sale price, star rating and review count, Bestseller /
New / Sephora Exclusive / Limited Edition / Online Only flags, number of
variants, image and product URL.

Give it category slugs or URLs (`face-serum`, `lipstick`,
`https://www.sephora.com/shop/moisturizing-cream-oils-mists`); it reads every
page of each category, 60 products per page.

HTTP only, no login, no browser.

### What it is for

- **Price monitoring** — list and sale prices across a whole category.
- **Assortment and brand research** — who sells what in a category, which
  products are bestsellers, new, or Sephora exclusives.
- **Review-volume benchmarking** — rating and review count for every product.

### Input

| field | what it does |
| --- | --- |
| `categories` | Sephora `/shop/` category slugs or URLs. One search each. |
| `maxItems` | Products per category (default 120, `0` = the whole category). |
| `minRequestInterval` | Seconds between requests. Never below 5 (see below). |
| `proxyConfiguration` | US residential by default (see below). |

Search (`/search?keyword=`) and brand pages are not supported: Sephora's
robots.txt disallows search, and this Actor only reads the category pages it
was measured against. An unsupported URL gets an `invalid_input` row.

### Things about this site worth knowing before you trust a run

#### 1. Prices are often ranges, and a sale can cover only some sizes

`listPrice` is `"$82.00 - $119.00"` for a product sold in several sizes
(about a quarter of products). The Actor adds `priceMin` / `priceMax`.
A discount adds `salePrice`, and `saleScope` says whether it covers
`all_variants` or `some_variants`. A partial sale can equal the bottom of the
list range — `"$37.00 - $39.00"` with a `$37.00` sale is the $39 size marked
down — so compare against `priceMax`, not `priceMin`, in that case.

#### 2. Sorting is not available

Asking for the site's own "Bestselling" order by URL is refused by Sephora's
firewall (HTTP 403). Rows come in Sephora's default order, with `searchRank`.
Bestsellers are still flagged per product (`isBestseller`).

#### 3. The last page is empty, not an error

Face Serums reports 517 products: pages 1–8 hold 60, page 9 holds 37, and
page 10 answers normally with nothing on it. The Actor stops at the site's
own product count and checks every page it gets is the page it asked for.

#### 4. Delivery flags are left out

Sephora's listing marks every product "not eligible" for pickup, same-day
and even ship-to-home when no store is chosen. Those five fields describe a
store the page doesn't have, so they are dropped rather than passed on.

#### 5. Expect refusals — they are retried, and counted

Sephora's firewall (Akamai) judges each **address + browser fingerprint
pair**: the same fingerprint passes from one IP and is refused from the
next. No single fingerprint or network was reliable in testing, so every
refusal moves to another browser fingerprint and a fresh US residential IP
(up to 10 tries per page). Measured: a full 517-product category took 12
requests with 2 refused; a bad minute refused 19 of 28. Each summary row
reports `refusedRequests` and `totalRequests`.

If a page still can't be read, the category **keeps the products already
read**: its summary says `stoppedBecause: "refused_at_page_N"` and an
`ERROR` row (`fetch_failed_partial`) explains. Run again for the rest.

A page of 60 products is 200–400 KB of residential traffic; the 517-product
category cost $0.02 in total.

Sephora's robots.txt asks crawlers to wait 5 seconds between requests. This
Actor always does, across all categories in a run — a 500-product category
takes about 45 seconds.

### Output

One `PRODUCT` row per product — Sephora's own listing object with parsed
fields added — plus one `CATEGORY_SUMMARY` per category and an `ERROR` row
for anything that could not be read.

```json
{
  "recordType": "PRODUCT",
  "brandName": "MAC Cosmetics",
  "displayName": "M·A·Cximal Silky Matte 12HR Wear Lipstick",
  "priceMin": 16.0,
  "priceMax": 25.0,
  "priceIsRange": true,
  "onSale": false,
  "ratingValue": 4.5929,
  "reviewCount": 759,
  "isBestseller": true,
  "productId": "P510799",
  "skuId": "2837425",
  "productUrl": "https://www.sephora.com/product/mac-cosmetics-m-a-cximal-silky-matte-lipstick-P510799?skuId=2837425",
  "categoryName": "Lipstick",
  "categoryPath": ["Makeup", "Lip", "Lipstick"],
  "searchPage": 1,
  "searchRank": 1
}
```

```json
{
  "recordType": "CATEGORY_SUMMARY",
  "_input": "face-serum",
  "categoryName": "Face Serums",
  "categoryPath": ["Skincare", "Treatments", "Face Serums"],
  "siteTotal": 517,
  "pagesFetched": 2,
  "returnedCount": 70,
  "stoppedBecause": "maxItems",
  "refusedRequests": 0,
  "totalRequests": 2
}
```

### Limits

- Sephora US (sephora.com) only, prices in USD.
- Product pages (ingredients, full review text, every shade) are not fetched;
  each row links to its product page.

# Actor input Schema

## `categories` (type: `array`):

Sephora US category slugs or URLs, one search each: 'face-serum', 'lipstick', 'moisturizing-cream-oils-mists', or https://www.sephora.com/shop/face-serum. Only /shop/ category pages are supported (Sephora's robots.txt disallows search).

## `maxItems` (type: `integer`):

0 = the whole category. 60 products per page; pages are 5 s apart (robots.txt Crawl-delay), so 500 products take about 45 s.

## `minRequestInterval` (type: `number`):

Never below 5 -- Sephora's robots.txt asks for a 5-second crawl delay, and a lower value is raised to it.

## `proxyConfiguration` (type: `object`):

US residential by default. Sephora's Akamai judges each address + browser fingerprint pair; Apify's own IPs and the datacenter pool were refused far more often than US residential. 200-400 KB of residential traffic per page of 60 products (measured).

## Actor input object example

```json
{
  "categories": [
    "face-serum"
  ],
  "maxItems": 120,
  "minRequestInterval": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "categories": [
        "face-serum"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/sephora-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "categories": ["face-serum"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/sephora-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "categories": [
    "face-serum"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call scrapyx/sephora-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/sephora-products-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MTeNKLnzEuLVnOPnb/builds/4Rvj6YOPM3xDYufCj/openapi.json
