# Olive Young Product Scraper: Ingredients & Prices (`pradio/olive-young-product-scraper`) Actor

Scrape Olive Young products from its global store by keyword, category, product URL or best-seller page: price, discount, rating, stock and full ingredient lists, volume and manufacturer on every product at the base price. Optional review statistics. No login.

- **URL**: https://apify.com/pradio/olive-young-product-scraper.md
- **Developed by:** [Pradio Actors](https://apify.com/pradio) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.67 / 1,000 product returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Olive Young Product Scraper: Ingredients & Prices

### What does Olive Young Product Scraper do?

Olive Young Product Scraper reads product listings from Olive Young's global store and returns one row per product. It returns full ingredient lists, volume and manufacturer on every product at the base price. Feed it a keyword, a category URL, a product URL, a product number or the store's best-seller page. Up to 1,000 rows land in your dataset on a run, each carrying the 46 fields the example below shows.

On a measured run over 14 brand names this Actor had never seen, 10 answered with 92 rows in total. The other 4 were names the store does not carry, and each came back as an uncharged status row. You pay $0.004 per product row, only for rows whose status is ok. Paid plans pay less: down to $0.00067 a row on Gold and above.

### Who uses Olive Young Product Scraper

| Buyer | What they run it for |
|---|---|
| K-beauty resellers | Price, discount and stock checks across the catalogue before listing |
| Brand and category analysts | Tracking ratings, review counts and promotions by brand |
| Trend researchers | Watching which products Olive Young flags best, trending or new |
| Sourcing and compliance teams | Pulling a category's shelf into rows with ingredients and manufacturer |

### Features

- **Real keyword search.** Your keyword is answered by Olive Young's own search index, the same index the store's search page calls. A query for "laneige" returns LANEIGE products, not the unfiltered catalogue. The store's index decides relevance, so a broad term like "cream" can return the adjacent products the site itself shows, exactly as it shows them.
- **Category URLs with the store's own filters.** Filter on Olive Young's own page first, then paste the filtered URL. Every product on the page comes back as a row, including the stock counts, attribute labels and spec groups a keyword search does not publish.
- **Product URLs and product numbers.** Paste a detail page URL or a bare product number (GA… or GS…) and get exactly that product's row.
- **Full ingredient lists, volume and manufacturer on every product at the base price.** Each row reads the product's notice panel.
- The notice panel carries its ingredient names as an array, content volume or weight, skin suitability, shelf life, manufacturer and distributor, country of manufacture and the store's functional-cosmetic flag. `includeProductInfo` is on by default; turning it off skips that read. How-to-use, precautions and the service number are never emitted.
- **The store's best-seller page.** Paste `https://global.oliveyoung.com/display/page/best-seller` and the store's ranked list comes back as rows. `position` is the product's rank in it.
- **Price and rating filters plus a sort dial.** `minPrice`, `maxPrice` and `minRating` drop rows before they are emitted, so a filtered-out product is never billed. `sort` orders the results: the store sorts its own list on a category URL; on other seeds the collected rows are re-ordered before they are returned.
- **Several seeds in one run.** Put one keyword or URL per line in `keyword` and each is read separately. Every row echoes the line it came from in `input`.
- **Optional review statistics.** With `includeReviews` on, each product row carries aggregate review figures: average rating, 1-to-5-star counts, verified-buyer share, the latest review's date, and skin-type counts when at least 5 reviews exist. Aggregates only: no review text, no reviewer data, ever.
- **One row per product, charged after it lands.** Duplicates across overlapping seeds are folded before the charge. You are billed only for rows in your dataset whose status is ok.

### What you can count on

- You pay only for rows whose status is ok. An entry the run cannot answer is pushed as an uncharged `ITEM_STATUS` row that names the entry and says why.
- Every row is charged only after it is written to your dataset. A row you cannot see is never billed.
- A run that finds nothing returns one uncharged status row that says so, never an empty dataset you have to second-guess. This source always reports the miss as an `ITEM_STATUS` row, so the contract's reserved `NO_MATCHING_ITEMS` type never fires here.
- A spending limit stops the run cleanly. The `STOPPED_EARLY` row says how many rows were returned and how many were not.
- Every run writes a RUN\_SUMMARY with rowsFetched, rowsPushed, rowsCharged and duplicatesDropped, so a short run and a broken one are told apart.
- A refusal (a block, a CAPTCHA, a 401/403/429) stops the run at that entry, which becomes an uncharged status row naming what happened; the rows already read stand. An error that is not a refusal fails the run with the error in the log. It never returns rows full of nulls and calls it success.
- No value is invented. A field the store does not publish for that product is null, and the output table below says which.

### Why this one

The most-used alternative on this platform sends your keyword to a list call that ignores it. On 2026-09-26 its "laneige" run and a nonsense term returned the same unfiltered catalogue, the same 10,000 hits. The term you type changed nothing it sent back. This Actor calls the store's own search instead, so the term you type is the term that is answered.

Run on the same "cream" seed on 2026-09-27, it returned 20 rows, its own cap; this Actor returned 100. And on every one of those 20 rows it left the fields a product page exists for empty. The ingredient list, the volume, the country of manufacture, the per-option prices and barcodes, the promo label, the trending and selling-fast flags. This Actor filled them on almost every row of the same run: ingredients on 95 of 100, options and product barcodes on 98, country of manufacture on 96, the promotion label on 81, the trending flag on 65.

It is also the cheaper read. The alternative charges $0.0078 a product row plus a $0.05 fee on every run, and bills detail extraction as a separate event on top. This Actor charges $0.004 a row plus the platform's own start event, $0.00005 a run at the default memory. The notice panel is included at the base price.

A miss costs nothing on this Actor, and the row it leaves says why.

### What data does Olive Young Product Scraper return?

One row per product. The row below is the shape a default run returns: 46 fields, every one declared in the dataset schema. A field the store does not publish for a product arrives as null.

| Field | Meaning |
|---|---|
| `url` | The product's detail page URL |
| `product_title` | The product's display title as the store lists it |
| `brand` / `brand_id` | Brand name and Olive Young's brand identifier |
| `price` | Selling price as a number, in `currency` |
| `original_price` | List price before discount, when the store shows one |
| `discount_rate` | Percentage off the store shows. The store can show a rate without publishing a list price; `original_price` is then null |
| `option_price_min` / `option_price_max` | Lowest and highest option prices the store lists |
| `currency` | Price currency code (USD on the global store) |
| `on_promotion` | The store's promotion flag |
| `inStock` / `sold_out` | Whether it is purchasable now, and the store's sold-out flag |
| `stock_qty` | Sellable stock count; published only on the category list |
| `displayed_rating` / `review_count` | The star rating the product displays, and the store's review count |
| `category` / `category_id` / `category_path` | Leaf category, its identifier, and the full breadcrumb |
| `attributes` | The store's attribute labels, for example Brightening, Vegan; null on every read that is not a category page, where the store publishes them |
| `specs` | Grouped spec values the store publishes: skin concern, formulation, featured ingredients; null on every read that is not a category page |
| `thumbnailImage` / `images` / `image_count` | Links to the product's images on Olive Young's CDN, and how many links the row carries |
| `product_id` | Olive Young's product number (GA… or GS…) |
| `gtin` | The product's barcode (EAN/GTIN) as the store records it; filled on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on; null on category rows (and best-seller rows with it off) |
| `options` | Every purchasable variant, each an object with its name, price, list price, sold-out state, own barcode and image link; filled on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on |
| `subtitle` | The product's subtitle line (editions, shipping notes); filled on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on |
| `selling_fast` / `is_trending` | The store's selling-fast and trending flags; filled on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on |
| `promo_label` | The promotion's label text as the store prints it, for example SALE; filled on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on |
| `is_new` / `is_best` | The store's new and best-seller flags; they fill on detail reads and the best-seller page, and `is_new` is the store's own flag, set rarely |
| `product_details` | Extra store facts: product type, original Korean name, has-options, release date, review notes |
| `position` | The product's position within its result page; on the best-seller page, its rank |
| `input` | The input line this row came from, echoed back so several seeds in one run stay attributed |
| `status` | `ok` on a data row |

With `includeReviews` on, each product row also carries review aggregates:

| Field | Meaning |
|---|---|
| `rating_avg` | Average rating computed over the reviews this run read, where `displayed_rating` is the page's own figure |
| `rating_distribution` | Counts of 1-to-5-star reviews, for example `{"1": 12, "2": 5, "3": 30, "4": 210, "5": 989}` |
| `verified_share` | Share (0 to 1) of read reviews the store marks as verified purchases |
| `latest_review_date` | Date of the product's newest review |
| `skin_type_distribution` | Counts of reviewer skin types, for example `{"Dry": 8, "Oily": 21, "Other": 6}`. Withheld under 5 reviews; buckets under 5 fold into Other |
| `reviews_covered` | How many reviews the aggregates were computed over; equals `review_count` when all were read |

With `includeProductInfo` on (the default), each product row also carries its notice panel:

| Field | Meaning |
|---|---|
| `volume` | Content volume or weight, for example `60mL*2ea` |
| `suitable_for` | The skin types or users the panel recommends the product for |
| `shelf_life` | Shelf life and period-after-opening as the panel states them |
| `manufacturer` | Manufacturer and responsible distributor |
| `country_of_manufacture` | Country of manufacture |
| `ingredients` | Ingredient names as an array, split on commas with the store's variant tags stripped |
| `ingredients_text` | The panel's ingredient block as the store prints it, whitespace-normalised, variant labels kept |
| `functional_cosmetic` | Whether the panel flags the product a functional cosmetic under the Korean regulator's rules |

The panel's how-to-use, precautions, service number and compensation boilerplate are read but never emitted.

A status or run-level row instead carries:

| Field | Meaning |
|---|---|
| `row_type` | `ITEM_STATUS` on an entry the run could not answer; its `status` field names the verdict. `NO_MATCHING_ITEMS` is the contract's type for a run that pushed no rows at all; this source always reports the miss as an `ITEM_STATUS` row instead. `STOPPED_EARLY` when the run spend limit ended it. Data rows carry `ROW` |
| `status` | On an `ITEM_STATUS` row, the miss verdict: `not_found`, `bad_url`, `blocked` and so on |
| `reason` | Why the entry was not answered, why the run returned no rows, or why it stopped early |
| `rowsFetched` | How many rows the source handed over before de-duplication and the cap |
| `rowsReturned` | How many data rows are in the dataset (0 on an empty result; the count before the run spend limit on STOPPED\_EARLY) |
| `rowsRemaining` | How many fetched rows were not returned (0 on an empty result) |

#### A real row

This is a captured row from the prefilled `cream` run, not a typed example. The term is broad, so the run's other rows include the adjacent products the store's own index shows shoppers: serums and ampoules beside the creams, exactly as the site returns them:

```json
{
  "thumbnailImage": "https://cdn-image.oliveyoung.com/uploads/images/display/prdtImg/1149/b2e09d76-bf30-4b8b-9c1a-ae619d523883.jpg",
  "url": "https://global.oliveyoung.com/product/detail?prdtNo=GA231121124",
  "brand": "S.NATURE",
  "price": 25.95,
  "currency": "USD",
  "inStock": true,
  "category": "Cream",
  "product_id": "GA231121124",
  "product_title": "[Sanrio EDITION] S.NATURE Aqua Squalane Moisturizing Cream 60ml+60ml (2 Options)",
  "input": "cream",
  "status": "ok",
  "brand_id": "B00436",
  "original_price": 50,
  "discount_rate": 48.1,
  "on_promotion": true,
  "option_price_min": 25.95,
  "option_price_max": 25.95,
  "displayed_rating": 4.8,
  "review_count": 1545,
  "stock_qty": null,
  "sold_out": false,
  "is_new": false,
  "is_best": true,
  "category_path": ["OliveYoungGlobal", "Skincare", "Moisturizers", "Cream"],
  "category_id": "1000000016",
  "attributes": null,
  "specs": null,
  "images": ["https://cdn-image.oliveyoung.com/uploads/images/display/prdtImg/1149/b2e09d76-bf30-4b8b-9c1a-ae619d523883.jpg","https://cdn-image.oliveyoung.com/uploads/images/display/prdtImg/1801/a0035f21-a707-436e-8d8d-af000ac40035.jpg","https://cdn-image.oliveyoung.com/uploads/images/display/prdtImg/1252/7c9eb53d-df42-48a1-a5b4-8e1d7d091139.png","https://cdn-image.oliveyoung.com/uploads/images/display/prdtImg/1760/03c7b36b-a2c5-4be9-8a8b-438b1eb532d5.jpg"],
  "image_count": 4,
  "product_details": {"product_type": "GENERAL_PRODUCT", "original_name": "에스네이처 아쿠아 스쿠알란 수분크림 60ml 더블 어워즈 한정기획 (2023)", "has_options": true, "released": null},
  "position": 1,
  "volume": "60mL*2ea",
  "suitable_for": "For all skin types",
  "shelf_life": "3 years after manufacture",
  "manufacturer": "Bio Costech Co., Ltd. / S. NATURE",
  "country_of_manufacture": "South Korea",
  "ingredients_text": "Water, Squalane (150,000ppm), Glycerin, 1,2-Hexanediol, Betaine, Panthenol, Sodium Hyaluronate, Hydrolyzed Collagen, Hydroxypropyltrimonium Hyaluronate, Hydrolyzed Hyaluronic Acid, Sodium Acetylated Hyaluronate, Acetyl Hexapeptide-8, Adenosine, Hyaluronic Acid, Hydrolyzed Sodium Hyaluronate, Sodium Hyaluronate Crosspolymer, Potassium Hyaluronate, Acrylates/C10-30 Alkyl Acrylate Crosspolymer, Arginine, Allantoin, Xylitylglucoside, Anhydroxylitol, Xylitol, Glucose, Butylene Glycol, Ammonium Acryloyldimethyltaurate/VP Copolymer, Caprylyl Glycol",
  "ingredients": [
    "Water", "Squalane (150,000ppm)", "Glycerin", "1,2-Hexanediol", "Betaine", "Panthenol", "Sodium Hyaluronate",
    "Hydrolyzed Collagen", "Hydroxypropyltrimonium Hyaluronate", "Hydrolyzed Hyaluronic Acid", "Sodium Acetylated Hyaluronate",
    "Acetyl Hexapeptide-8", "Adenosine", "Hyaluronic Acid", "Hydrolyzed Sodium Hyaluronate", "Sodium Hyaluronate Crosspolymer",
    "Potassium Hyaluronate", "Acrylates/C10-30 Alkyl Acrylate Crosspolymer", "Arginine", "Allantoin", "Xylitylglucoside",
    "Anhydroxylitol", "Xylitol", "Glucose", "Butylene Glycol", "Ammonium Acryloyldimethyltaurate/VP Copolymer", "Caprylyl Glycol"
  ],
  "functional_cosmetic": null,
  "gtin": "8809506312905",
  "options": [
    {"name": "[HELLO KITTY EDITION] Cream 60ml+60ml (+Keycap Keyring)", "price": 25.95, "original_price": 50, "sold_out": false, "gtin": "8809506312905", "image": null},
    {"name": "[SET] Cream 60ml+60ml", "price": 25.95, "original_price": 50, "sold_out": false, "gtin": "8809506311052", "image": null}
  ],
  "subtitle": null,
  "selling_fast": true,
  "is_trending": null,
  "promo_label": "SALE",
  "row_type": "ROW"
}
```

### How much does it cost?

Each product row written to your dataset bills one `item-returned` event: **$0.004 on the Free plan**, charged only after the row is in your dataset. Higher-volume plans pay less per row: $0.00133 on Bronze, $0.001 on Silver, and $0.00067 on Gold, Platinum and Diamond. Apify also bills its own `apify-actor-start` event on every run: $0.00005 per event, one per GB of run memory, minimum one. At this Actor's memory allocation it is charged once a run, and this Actor adds nothing on top.

Status rows, filtered-out products, misses and dropped duplicates are never charged, and neither is the RUN\_SUMMARY.

| Rows returned | You pay |
|---|---|
| 100 products | $0.40 + $0.00005 start |
| 1,000 products | $4.00 + $0.00005 start |
| 10,000 products, over ten runs (one run is hard-capped at 1,000) | $40.00 + ten start fees |

In the buyer's unit: at the measured answer rate, a run of 100 brand searches returns about 650 product rows for about $2.60. The same at 1,000 searches is about $26, across the several runs the 1,000-row cap spreads them over. A search that finds nothing bills no product row.

`includeProductInfo` reads the notice panel once per product and `includeReviews` reads its review pages, so a run with both on takes longer. Neither changes the per-row price.

### How do I use Olive Young Product Scraper?

1. Open the Actor and press Try for free.
2. Set `keyword` to a search term like `cream`, a category URL like `https://global.oliveyoung.com/display/category?ctgrNo=1000000008`, the best-seller page `https://global.oliveyoung.com/display/page/best-seller`, a product URL or product number, or several of those one per line.
3. Press Start. Rows stream into the dataset as they are read. Export them as JSON, CSV or Excel, or pull them over the API.

#### Worked examples

Use it to track products by keyword, review aggregates included:

```json
{
  "keyword": "laneige\ncosrx",
  "maxItems": 100,
  "includeReviews": true
}
```

Use it to read one category shelf end to end:

```json
{
  "keyword": "https://global.oliveyoung.com/display/category?ctgrNo=1000000008",
  "maxItems": 200,
  "includeReviews": false
}
```

Use it to pull the best-seller page in rank order, keeping only products under a maximum price you set:

```json
{
  "keyword": "https://global.oliveyoung.com/display/page/best-seller",
  "maxItems": 50,
  "maxPrice": 30
}
```

Or call it over the API:

```bash
curl -X POST "https://api.apify.com/v2/acts/Pradio~olive-young-product-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" -H "Content-Type: application/json" -d "{\"keyword\": \"laneige\", \"maxItems\": 100}"
```

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `keyword` | string | prefilled `cream` | One seed per line: a search term, a category page URL, a product page URL or a product number (GA… or GS…). Each row echoes its line in `input` |
| `maxItems` | integer | 100 | The most product rows one run returns, hard-capped at 1,000. The run stops at the cap and the run log says so |
| `includeReviews` | boolean | false | Also read each product's review aggregates: average rating, star counts, verified-buyer share, latest review date, skin-type counts. Extra requests per product; aggregates only |
| `includeProductInfo` | boolean | true | Also read each product's notice panel: ingredient list, volume, skin suitability, shelf life, manufacturer and distributor, country of manufacture, functional-cosmetic flag. One extra read per product at the same per-row price |
| `minPrice` | number | none | Keep only rows priced at or above this USD figure; a dropped row is never billed |
| `maxPrice` | number | none | Keep only rows priced at or below this USD figure; a dropped row is never billed |
| `minRating` | number | none | Keep only rows whose displayed rating is at least this (0 to 5); unrated products are dropped too, and a dropped row is never billed |
| `sort` | select | `popular` | `popular`, `newest`, `price_low`, `price_high`, `most_reviews` or `highest_rated` |

#### Where filters and sort are applied

On a category URL the store does the work itself: `minPrice` and `maxPrice` go out as its own `lowestPrice` and `highestPrice` parameters, `sort` as its own `prdtSortStdrCode`. The emitted rows are still re-checked before they are returned. On keyword, product-URL, product-number and best-seller seeds the store takes no filter or order, so every bound is a local check on the collected rows before emit. `minRating` is always a local check, on every seed kind: the store's own rating facet is a band picker, not an exact floor. Either way the rule is the same: a row outside your filters is dropped before it is emitted and never billed.

### Output

Rows land in the default dataset as they are read. The Overview view shows title, brand, price, category and stock at a glance; the All fields view carries every declared field, including the ones only some seed kinds fill.

How a run ends is visible in the data. A normal run ends in data rows whose `row_type` is `ROW`. An entry the run cannot answer becomes an uncharged `ITEM_STATUS` row naming it: a `bad_url`, a product number or term with no answer (`not_found`), a `blocked` read. A run that finds nothing at all still pushes that one status row, never an empty dataset. `NO_MATCHING_ITEMS` is the contract's reserved type for a run that pushed no rows of any kind; this source always reports the miss as an `ITEM_STATUS` row instead. And a run stopped by the spend limit you set on the platform pushes one `STOPPED_EARLY` row carrying `rowsReturned` and `rowsRemaining`, also uncharged.

Every run also writes a RUN\_SUMMARY in the key-value store: rowsFetched, rowsPushed, rowsCharged and duplicatesDropped. Rows are de-duplicated by `url` before anything is pushed or charged, so overlapping seeds do not double-bill. A zero result is therefore an answer with counts behind it, never silence.

### What can you do with the data?

A reseller repricing a shop schedules a keyword run for its brands each morning and joins rows to its own listings on `product_id`. Price, discount rate, stock flags and the per-option prices tell it what moved overnight.

A category analyst tracks a rival's shelf. The same category URL weekly gives rating, review count, position and promotion flags per product; `includeReviews` adds verified share and skin-type counts.

A sourcing team pulls a category end to end before a buying trip. Ingredients as an array, manufacturer, country of manufacture and the functional-cosmetic flag make the shortlist checkable without opening hundreds of product pages.

A trend watcher reads the best-seller page on a schedule. `position` is the store's own rank; `selling_fast`, `is_trending`, `is_best` and `promo_label` show what the store is pushing this week.

Integrations are built in: Zapier, Make, webhooks and any tool that reads Apify datasets all work, and the API returns rows directly.

### Use Olive Young Product Scraper with AI agents

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=Pradio/olive-young-product-scraper"
```

Paste that line and your agent can call this Actor as a tool, pass it a keyword, and work with the rows it returns.

### Personal data

This Actor returns product facts and review statistics only. When `includeReviews` is on, review information arrives solely as per-product aggregates: rating averages, star counts, verified share, the newest review's date and skin-type counts. It never returns review text, excerpts, reviewer names, handles or IDs. Skin-type figures are withheld when fewer than 5 reviews exist.

The lawful basis for the processing behind these rows is legitimate interest in indexing publicly available product information on Olive Young's global store (global.oliveyoung.com), read logged out. No personal identifier present in a raw response is ever persisted into your dataset, retained beyond the run, or shared. Any incidental personal data in a source response is dropped before the row is written and is deleted on request. Pradio is the controller for this Actor's processing. For any erasure or access request, contact me through the Store's Issues tab or the publisher profile.

### Is it legal to read Olive Young?

Only you can decide how you may use the rows you get. That responsibility is yours, so read the site's terms and your own obligations first. Olive Young's terms restrict commercial reuse of its information; treat the rows accordingly. What this Actor does is narrow on purpose. It sends logged-out reads, one page per request, at a slow pace. It returns facts about the products on that page: name, price, brand, rating, stock flags. It does not take a bulk copy of the site, and it never tries to get past a protection. A block, a CAPTCHA or an objection stops the run for that entry. The status row says so, uncharged.

A note on search, honestly. Olive Young's search page refuses automated reads, so this Actor does not fetch it. Keyword input is answered by the store's own product-data call, the same one its pages make, measured open on 2026-09-26. If it ever refuses, the run stops and says so rather than working around it.

A site owner who objects can open an issue on this page's Issues tab or contact me through the Store. The site is removed from the Actor the same day.

### Transparency

For transparency, this is exactly what the Actor does. It reads the public product pages of Olive Young's global store (global.oliveyoung.com), logged out, with no account, one request per page: the same product-data calls the store's own pages make. The facts it collects are the ones a shopper sees. Product names, brands, prices, discounts, ratings, review counts, stock and promotion flags, categories, image links, product codes, option lists, and each product's notice panel (ingredient list, volume, shelf life, manufacturer and distributor, country of manufacture, functional-cosmetic flag). With `includeReviews` on it also reads review pages and returns only aggregate review statistics, never review text or reviewer data. Every field it returns is named in the output table above, and the store's own pages show the same facts to any visitor. Olive Young's terms restrict commercial reuse of site information; treat the rows accordingly.

If you own a page these rows describe, or spot a fact that is wrong, ask for a correction or deletion through the Store's Issues tab or the publisher profile. It is handled the same day.

### Release notes

- `0.1`: first public build. Product listings by keyword, category URL, product URL, product number or the store's best-seller page, one row per product. Full ingredient lists, volume and manufacturer on every product at the base price behind `includeProductInfo` (on by default); `minPrice`, `maxPrice` and `minRating` filters applied before the charge, and a `sort` dial; per-product review aggregates behind `includeReviews`. Per-row pricing with every miss uncharged, and `maxItems` hard-capped at 1,000 rows a run.

### Limits

- `stock_qty`, `attributes` and `specs` fill only on category-seeded rows. The store publishes them only on its category list; keyword-search and product-URL reads do not carry them, so all three are null there.
- `gtin`, `options`, `subtitle`, `selling_fast`, `is_trending` and `promo_label` fill on keyword, product-URL and product-number seeds, and on best-seller rows while `includeProductInfo` is on. On category rows (and best-seller rows with it off) they are null.
- `original_price` is null whenever the store shows no strikethrough list price; it was filled on 63% of the measured sample (58 of 92 rows).
- `is_new` is the store's own flag and the store sets it rarely: a measured pass over its own new-arrivals page on 2026-09-28 found almost no products carrying it. A row of `false` is the store's answer, not a gap.
- Review output is aggregates only, by design and forever: no review text, no reviewer names, handles or IDs. `skin_type_distribution` stays null under 5 reviews, and buckets under 5 fold into Other.
- The run reads one page per request at a fixed pace. A 100-row run takes a few minutes. There is no fast mode, by design.
- A run is bounded by `maxItems`, hard-capped at 1,000 rows, and by the run spend limit you set on the platform (maxTotalChargeUsd). The spend limit ends the run with an uncharged `STOPPED_EARLY` row; the cap is named in the run log and in RUN\_SUMMARY's counts. This Actor never bulk-copies the catalogue.
- A keyword the store does not carry adds no rows and bills none. A run that finds nothing at all returns one uncharged row whose `status` is `not_found`. You never get zero rows and silence.

### Troubleshooting

**Fewer rows than `maxItems`?** The store ran out of matches for your seed. RUN\_SUMMARY's `rowsFetched` shows what it actually had, and the run log names the cap when `maxItems` is what stopped it.

**Only one row in the dataset?** Read `row_type` and `reason`. `ITEM_STATUS` means an entry could not be answered; the row's `status` names the verdict: `bad_url` for a malformed or off-site link, `not_found` for a product number or term the store does not answer, `blocked` when the site refused the read. `NO_MATCHING_ITEMS` is the contract's type for a run with no rows at all; this source always reports the miss as an `ITEM_STATUS` row. Neither is charged.

**Nulls in a row?** The store does not publish that field for that product or on that kind of read. The output table names the three fields only category pages carry and the six only a detail read fills.

**The run failed?** A non-refusal error, a fetch failure or a changed page shape, fails the run with the error in the log. A refusal is different: it stops the run at that entry as an uncharged `blocked` status row, and the rows already read stand. The Actor never papers over a refusal with nulls.

### FAQ

**Can I use integrations with Olive Young Product Scraper?**
Yes. It writes to an Apify dataset, so Zapier, Make, webhooks, schedules and every tool that reads Apify datasets work without extra setup.

**Can I use Olive Young Product Scraper with the Apify API?**
Yes. The input is plain JSON and the run-sync endpoint returns the dataset rows directly; the curl line above is a working example.

**Can I use it through an MCP server?**
Yes. Paste the `claude mcp add` line from the agents section and your agent gets this Actor as a tool, with the keyword input it expects.

**Is it legal to scrape Olive Young?**
See "Is it legal to read Olive Young?" above. The short version: public product pages read logged out, one request per page, facts only, and it stops the moment the site objects. You are responsible for having the right to read the pages you give it and for how you use the rows.

**How am I charged?**
$0.004 per product row whose status is ok on the Free plan, down to $0.00067 on Gold and above. The platform's $0.00005 start event per GB of run memory comes on top. Filtered-out rows, misses, status rows and duplicates bill nothing.

**Do I need proxies, and what happens when the site blocks a request?**
There is no proxy to configure, and one is not a way past a refusal here. If a page answers a block, a CAPTCHA or an objection, the run stops for that entry and the status row says so, uncharged. The fix is to wait and re-run, never to work around the refusal.

**How many products can one run return?**
`maxItems` sets it, default 100, hard-capped at 1,000 rows. Reads are paced one per page; a run is the seeds you name, never a bulk copy of the site.

**What output formats are available?**
Rows land in the run's dataset as JSON and can be exported as CSV, Excel, XML, RSS or HTML from the dataset view, or read directly through the Apify API. Every field in the output table is in each export.

**How fresh is the data?**
Every run reads the store's live pages at run time and nothing is cached between runs. A row is what the store published in the seconds that request took: price, stock, rating and promotion flags as they stood then.

### Feedback

Found a bug or a field you need? Open an issue on this page's Issues tab; it is the fastest way, and it is answered within two days. If this Actor saved you time, a review on the Store helps other buyers find it.

### Not affiliated

This Actor is an independent product. It is not affiliated with, endorsed by or sponsored by Olive Young or CJ. It reads only public, logged-out pages on global.oliveyoung.com. Olive Young's own terms restrict commercial reuse of its information, and how you use returned rows is your responsibility. No Olive Young logos or marks are used.

# Actor input Schema

## `keyword` (type: `string`):

A search term (for example 'laneige'), a category page URL, a product page URL or a product number (GA… or GS…) on Olive Young's global store. One seed per line runs several; each row echoes the line it came from in `input`.

## `maxItems` (type: `integer`):

The most product rows one run returns; the run stops at the cap and the run log says so.

## `includeReviews` (type: `boolean`):

Also read each product's review aggregates: average rating, star counts, verified-buyer share, latest review date, skin-type counts. Costs extra requests per product. Aggregates only: never review text or reviewer data.

## `includeProductInfo` (type: `boolean`):

Also read each product's notice panel: the ingredient list, volume, skin suitability, shelf life, manufacturer and distributor, country of manufacture, and whether the store flags it a functional cosmetic. One extra read per product at the same per-row price.

## `minPrice` (type: `number`):

Keep only rows priced at or above this USD figure. Rows outside the range are dropped before they are emitted, so they are never billed.

## `maxPrice` (type: `number`):

Keep only rows priced at or below this USD figure. Rows outside the range are dropped before they are emitted, so they are never billed.

## `minRating` (type: `number`):

Keep only rows whose displayed rating is at least this (0–5). Rows below it, and rows the store shows no rating for, are dropped before they are emitted, so they are never billed.

## `sort` (type: `string`):

The order rows come back in. On a category page URL the store sorts its own list; on a keyword, product URL or product number the store's route takes no order, so the collected rows are re-ordered before they are returned. On the best-seller page the store's rank order always stands and position is the rank.

## Actor input object example

```json
{
  "keyword": "cream",
  "maxItems": 100,
  "includeReviews": false,
  "includeProductInfo": true,
  "sort": "popular"
}
```

# Actor output Schema

## `rows` (type: `string`):

The product rows for this run: one row per listing read from Olive Young.

## `summary` (type: `string`):

Counts for this run: rows fetched, pushed, charged, duplicates dropped, whether it stopped early.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keyword": "cream"
};

// Run the Actor and wait for it to finish
const run = await client.actor("pradio/olive-young-product-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "keyword": "cream" }

# Run the Actor and wait for it to finish
run = client.actor("pradio/olive-young-product-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keyword": "cream"
}' |
apify call pradio/olive-young-product-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pradio/olive-young-product-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lN4aZIrm9kHeuxtNr/builds/xfJZHnaIUBAHvCgfx/openapi.json
