# Inditex Product Scraper — 7 Brands, One Dataset (`sian.agency/inditex-product-scraper`) Actor

Seven Inditex brands in one dataset: Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho and Pull\&Bear, all 58 columns, so rows from different brands compare directly. Price, sizes, stock, barcodes across 200+ country stores. No account or proxy needed.

- **URL**: https://apify.com/sian.agency/inditex-product-scraper.md
- **Developed by:** [SIÁN OÜ](https://apify.com/sian.agency) (community)
- **Categories:** E-commerce, Business
- **Stats:** 3 total users, 2 monthly users, 77.8% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.15 / 1,000 scraped products

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Inditex API — Zara + 6 Brands, One Dataset 🛍️

[![SIÁN Agency Store](https://img.shields.io/badge/Store-SI%C3%81N%20Agency-1AE392)](https://apify.com/sian.agency?fpr=sian) [![Myntra Product Scraper](https://img.shields.io/badge/Store-Myntra%20Product%20Scraper-FF3F6C)](https://apify.com/sian.agency/myntra-product-scraper?fpr=sian) [![Nike Product Scraper](https://img.shields.io/badge/Store-Nike%20Product%20Scraper-111111)](https://apify.com/sian.agency/nike-product-scraper?fpr=sian) [![AliExpress Product Scraper](https://img.shields.io/badge/Store-AliExpress%20Product%20Scraper-E62E04)](https://apify.com/sian.agency/aliexpress-product-scraper?fpr=sian)

#### 🎉 One Inditex API across Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho and Pull\&Bear — 58 identical columns per product, 220 country stores, every price in its own currency

##### The only place on the Apify Store where all seven brands land in one dataset with one schema. Everyone else sells them as seven separate rentals totalling $84.93 a month.

### 🔎 What is the Inditex API — and when should you use it?

The **Inditex Product Scraper** turns seven brand catalogues into one clean table. Pick the brands you want, pick a market, and every row comes back with the same 58 columns and a `brand` column to sort by. No account, no portal API key, no browser automation to maintain.

**Use it when you need:** cross-brand comparability. Price, old price, discount and currency, colours, images, taxonomy, descriptions, composition, care, and per-SKU sizes with stock and GTIN barcodes — in one schema across Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho and Pull\&Bear.

**Use something else when:** you only ever want one brand. The single-brand Actors are cheaper per row — [Zara Product Scraper](https://apify.com/sian.agency/zara-product-scraper?fpr=sian), [Zara Home Product Scraper](https://apify.com/sian.agency/zara-home-product-scraper?fpr=sian), [Bershka Product Scraper](https://apify.com/sian.agency/bershka-product-scraper?fpr=sian), [Massimo Dutti Product Scraper](https://apify.com/sian.agency/massimo-dutti-product-scraper?fpr=sian), [Stradivarius Product Scraper](https://apify.com/sian.agency/stradivarius-product-scraper?fpr=sian), [Oysho Product Scraper](https://apify.com/sian.agency/oysho-product-scraper?fpr=sian) and [Pull\&Bear Product Scraper](https://apify.com/sian.agency/pull-and-bear-product-scraper?fpr=sian).

### 🤖 Use with AI agents

Already connected to the [Apify MCP server](https://mcp.apify.com)? Just ask for this Actor by name: sian.agency/inditex-product-scraper

Otherwise copy this prompt into Claude, ChatGPT, Cursor or any MCP-enabled assistant:

```text
I want product data across several Inditex brands using the Apify Actor `sian.agency/inditex-product-scraper`.

Use it when I need: two or more of Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho and Pull&Bear in ONE dataset with one schema — price, old price, discount, currency, colours, images, taxonomy, description, composition, care, and per-SKU sizes with stock and GTIN barcodes.

Don't use it when: I only ever want one brand — the single-brand Actors are cheaper per row, e.g. `zara-product-scraper` or `bershka-product-scraper`.

How to call it: `brands` is the list of brands to run, one after another into the same dataset. `mode` is "overview" (walks each brand's category tree) or "detail" (reads product URLs I paste). Set `country` to the market, keep `allCategories` on, and cap the whole run with `maxItems` — it is a run total across all brands, not a per-brand cap.

Start with this input:
{
  "brands": ["zara", "bershka", "stradivarius"],
  "mode": "overview",
  "country": "de",
  "allCategories": true,
  "maxItems": 300
}

Ask me which brands and which market, then run the Actor and summarise the results as a table grouped by brand.
```

**Things you can ask your agent for:**

- *Compare average price per family across Zara, Bershka and Stradivarius in the Spanish market.*
- *Pull 200 products from each of the seven brands in Germany and show me who discounts hardest.*
- *Build me one table of all Inditex homeware and womenswear in France, with GTIN barcodes where they exist.*

Machine-readable API, MCP config and OpenAPI definition for this Actor are published at [apify.com/sian.agency/inditex-product-scraper.md](https://apify.com/sian.agency/inditex-product-scraper.md).

***

### 📋 Overview

**Seven catalogues, two very different upstream APIs, one table.** Zara and the six sibling brands publish product data in structurally different shapes — different ids, different price encodings, different bundle semantics. This Actor reads both and emits the same 58 columns either way, so a Bershka row and a Zara row sit side by side and mean the same thing.

**What you get:**

- 🏬 **Seven brands, one schema**: Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho, Pull\&Bear, each row stamped with its `brand`
- 🌍 **220 country stores**: the union across the seven brands, each priced in its own currency with the right divisor applied. The six shared-catalogue brands cover ~216 of them; zara.com is open in only 96, so check the Market field's note before pairing Zara with a small market
- 📐 **Per-SKU sizes and stock**: colour by size, with buyable and back-soon flags
- 🏷️ **GTIN barcodes**: the join key to every other retail dataset, on five of the six shared-catalogue brands — Massimo Dutti publishes none at all
- 🗓️ **Promotion windows**: per-SKU price start and end dates, so you know the date a markdown ends, not only that one is running
- 🧾 **A bill you can audit**: the `source` column on every row says which call produced it, so your invoice reconciles against the dataset
- 💰 **$3.50 per 1,000 products** on the category sweep, with no subscription. The only other vendor covering all seven brands charges $84.93 a month before you have scraped anything

### ✨ Features

- 🏬 **Pick your brands**: run one, three or all seven into a single dataset, sequentially
- 🗂️ **Category sweep**: walk each brand's whole tree, or name one category
- 🗺️ **Sitemap sweep**: pick up catalogue items no category lists, for genuine full-catalogue coverage
- 🔗 **Paste your own URLs**: URLs are routed to the brand their hostname names, so you can mix brands freely
- 🎚️ **Facet filters**: colour, size, category and discount shortcuts, or any facet group the site publishes
- 💶 **Price bounds**: keep only what falls between your floor and ceiling, in the market's major units
- 🧹 **Deduplicated**: the same product reached twice is counted and charged once, per brand
- 📉 **Honest run summary**: products pushed, duplicates skipped, ids that no longer exist, and fetches that genuinely failed — reported apart, never merged

### 🎬 Quick Start

Pick your brands, pick a market, cap the run, press Run.

```bash
curl -X POST 'https://api.apify.com/v2/acts/sian.agency~inditex-product-scraper/runs?token=YOUR_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
  "brands": ["zara", "bershka", "stradivarius"],
  "mode": "overview",
  "country": "de",
  "allCategories": true,
  "maxItems": 300
}'
```

### 🚀 Getting Started (3 Simple Steps)

#### Step 1: Choose your brands

**Brands** defaults to all seven. Brands run one after another into the same dataset, and every row carries a `brand` column, so a combined export stays sortable and groupable.

#### Step 2: Choose your market and scope

Set **Market** to a two-letter country. Leave **Sweep All Categories** on for the catalogue, or turn it off and name one category. Switch **Run Mode** to `detail` to read product URLs you paste instead.

#### Step 3: Cap it, run it, export it

**Max Products** is a run total across every brand you selected, not a per-brand cap — set it low for the first run. Export as JSON, CSV or Excel from the dataset tab.

**That's it. Within a few minutes you'll have:**

- One table covering every brand you picked, in one market
- Price, old price, discount and currency per product
- Per-SKU sizes with stock, and GTIN barcodes on five of the six shared-catalogue brands (not Massimo Dutti)
- A run summary telling you exactly what was and was not fetched

### 📥 Input Configuration

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| brands | array | No | Which brands to scrape: `zara`, `zarahome`, `bershka`, `massimodutti`, `stradivarius`, `oysho`, `pullandbear`. Defaults to all seven |
| mode | string | No | `overview` walks the categories, `detail` reads the URLs you paste. Default `overview` |
| country | string | No | Two-letter market, e.g. `DE`, `ES`, `FR`, `GB`, `JP`. The list is the union across the seven brands — ~216 markets on the shared-catalogue six, 96 on zara.com |
| language | string | No | Two-letter language for names and descriptions. Empty takes the market default |
| allCategories | boolean | No | Sweep every category. On by default, and it wins over Single Category |
| category | string | No | One category by slug or numeric id. Only used when the sweep is off |
| fromSitemap | boolean | No | Also crawl the product sitemap for anything the categories missed. Bills at the detail rate |
| maxItems | integer | No | Stop after this many distinct products — a run total across all brands. `0` means no cap. Default `100` |
| withSizes | boolean | No | Zara only: fetch per-SKU sizes, stock and the description. Charges both events on those rows |
| productUrls | array | No | Product URLs or bare ids, one per line. URLs route to the brand their hostname names |
| sort | string | No | Zara and the shared-catalogue brands use different sort ids. `default` is the only one that works for every brand at once |
| filters | array | No | Facet filters as `GROUP=VALUE`, one per line. Same group OR-ed, different groups AND-ed |
| minPrice | integer | No | Drop products below this price, in the market's major units |
| maxPrice | integer | No | Drop products above this price, in the market's major units |
| batchSize | integer | No | Product ids per call on the shared-catalogue brands. Default `100`, maximum `200` |
| jitterMs | integer | No | Random pause before each request, in milliseconds. Default `0` |
| proxyConfiguration | object | No | Leave it off. The Actor reaches these APIs directly and a proxy adds cost without adding success |

**Example — all seven brands, Spanish market:**

```json
{
  "brands": ["zara", "zarahome", "bershka", "massimodutti", "stradivarius", "oysho", "pullandbear"],
  "mode": "overview",
  "country": "ES",
  "allCategories": true,
  "maxItems": 1400
}
```

**Example — two brands, discounted items only:**

```json
{
  "brands": ["bershka", "stradivarius"],
  "mode": "overview",
  "country": "FR",
  "allCategories": true,
  "maxPrice": 40,
  "maxItems": 500
}
```

**Example — mixed-brand URLs:**

```json
{
  "mode": "detail",
  "country": "DE",
  "productUrls": [
    "https://www.bershka.com/de/t-shirt-c0p123456789.html",
    "https://www.oysho.com/de/leggings-c0p987654321.html"
  ]
}
```

### 📤 Output

Results are saved to the Apify dataset with **58 columns per product**, identical across all seven brands. Export as JSON, CSV or Excel.

| Field | Type | Description |
|-------|------|-------------|
| brand | string | Which brand the row belongs to: `zara`, `zarahome`, `bershka`, `massimodutti`, `stradivarius`, `oysho`, `pullandbear` |
| product\_id | integer/string | Catalogue id. The only column that is never null — the join key for everything else |
| url | string | Canonical product page for the market you scraped |
| source | string | Which call produced the row: `overview`, `detail`, or `overview+detail` |
| url\_id | integer/string | Numeric URL id on the six shared-catalogue brands. Null on Zara |
| reference | string | Internal style reference, stable across markets |
| display\_reference | string | The shorter reference printed on the page and the label |
| seo\_product\_id | string | Id used in Zara SEO URLs. Null on the shared-catalogue brands |
| seo\_keyword | string | URL slug for the product. Zara only |
| name | string | Product name in the market language |
| name\_en | string | English product name, when the catalogue carries one. Shared-catalogue brands only |
| product\_type | string | Catalogue type of the record, e.g. Product or Bundle |
| kind | string | Zara grid classification of the tile. Null on the shared-catalogue brands |
| country | string | Market the prices and names belong to |
| language | string | Language the names and descriptions came back in |
| store\_id | integer/string | Internal store id. Shared-catalogue brands only |
| catalog\_id | integer/string | Internal catalogue id. Shared-catalogue brands only |
| section\_name | string | Top level of the catalogue tree, e.g. WOMAN. Zara only |
| section\_name\_en | string | English name of the section. Shared-catalogue brands only |
| family\_name | string | Product family in the market language |
| family\_name\_en | string | English product family. Shared-catalogue brands only — group on `family_name` when Zara is in the run |
| subfamily\_name | string | Sub-level of the family in the market language |
| subfamily\_name\_en | string | English sub-level of the family. Shared-catalogue brands only |
| categories | array | Every category the product is filed under, as `{id, name}`. Shared-catalogue brands only |
| category\_id | integer/string | Id of the category this row was scraped from. Zara only |
| category\_name | string | Name of the category this row was scraped from. Zara only |
| price | number | Current price in major units, divisor already applied |
| old\_price | number | Pre-discount price. Null when the product is not reduced |
| currency | string | ISO currency of the market, e.g. EUR, GBP, JPY |
| discount\_pct | number | Percentage off, computed from price and old price |
| on\_special | boolean | Catalogue flag for a promotional product. Shared-catalogue brands only |
| main\_image | string | First image of the shown colour, full resolution |
| images | array | Every image URL for the shown colour, in catalogue order |
| description | string | Long product description. Zara grid rows carry none — turn on the size fetch |
| additional\_info | string | Extra copy the catalogue attaches to the product. Shared-catalogue brands only, and measured empty on every product we sampled |
| keywords | string | Merchandising keywords attached to the product. Shared-catalogue brands only, and measured empty on every product we sampled |
| assembly\_url | string | Link to an assembly or instruction sheet, mostly homeware and furniture. Shared-catalogue brands only |
| color\_name | string | Name of the colour this row represents |
| colors | array | Colour objects with id, name, reference and, on Zara, price, availability and hex |
| available\_color\_names | array | Colour names the grid offers for this style. Zara only |
| composition | array | Material breakdown per garment part, with the percentage of each fibre |
| care | array | Washing and care instructions as `{id, name, description}`. Shared-catalogue brands only — Zara's care text arrives inside `description` |
| variants | array | One entry per colour and size — see the table below |
| size\_guide | string | Set to `enabled` when the product page offers a size guide. Zara detail records only |
| availability | string | Stock state of the shown colour. Zara only — shared-catalogue brands report stock per SKU |
| is\_buyable | boolean | Whether the product can currently be added to a basket. Shared-catalogue brands only — Zara reports stock in `availability` |
| back\_soon | boolean | Catalogue flag for a restock that is already scheduled. Shared-catalogue brands only |
| visibility | string | Catalogue visibility state, e.g. `visible` or `hidden`. Shared-catalogue brands only |
| availability\_date | string | Date the product becomes or became available. Shared-catalogue brands only |
| first\_visible\_date | string | When the product first appeared in the catalogue. **Zara detail records only** |
| is\_continuity | boolean | Carryover line versus seasonal drop. Shared-catalogue brands only |
| is\_pinned | boolean | Whether merchandising pinned the tile to a fixed slot. Zara grid only |
| grid\_position | integer | Position of the tile inside the category grid. Zara grid only |
| join\_life | string | Join Life sustainability label text, when the product carries one. Shared-catalogue brands only |
| sustainability\_show | boolean | Whether the product page shows a sustainability badge. Shared-catalogue brands only |
| sustainability | object | Full sustainability node off the shown colour. Shared-catalogue brands only |
| traceability | object | Supply-chain traceability node. Shared-catalogue brands only |
| certified\_materials | array | Certified material entries off the shown colour. Shared-catalogue brands only |

**Inside `variants` — one entry per colour and size:**

| Field | Description |
|-------|-------------|
| sku | Stock-keeping unit id for this colour and size |
| color / color\_id | Colour name and id of the SKU |
| size / size\_id | Size label as the market prints it, plus Zara's size id |
| partnumber | Shared-catalogue part number |
| price / old\_price | SKU price and pre-discount price in major units |
| barcode | GTIN barcode of the SKU — shared-catalogue brands |
| price\_start\_date / price\_end\_date | When the current price took effect and when it expires |
| old\_price\_start\_date / old\_price\_end\_date | The same window for the pre-discount price |
| is\_buyable / back\_soon | Whether this SKU can be bought now, and whether a restock is scheduled |
| dimensions / weight | Named physical axes and gram weight of the SKU |
| origin | Manufacturing country of the SKU |
| availability / reference / demand | Zara per-SKU stock state, reference and sell-through signal |
| equivalent\_size\_id / twinned\_skus | Zara cross-market size id, and the same SKU under sibling style ids |

**Example row (trimmed):**

```json
{
  "brand": "bershka",
  "product_id": 123456789,
  "url": "https://www.bershka.com/de/t-shirt-c0p123456789.html",
  "source": "overview",
  "name": "OVERSIZE T-SHIRT",
  "price": 9.99,
  "old_price": 15.99,
  "currency": "EUR",
  "discount_pct": 37.5,
  "country": "DE",
  "family_name_en": "T-SHIRTS",
  "is_continuity": false,
  "variants": [
    { "sku": 44112233, "size": "M", "barcode": "8445123456789", "price": 9.99, "is_buyable": true,
      "price_start_date": "2026-08-01T00:00:00Z", "price_end_date": "2026-08-31T23:59:59Z", "weight": 180 }
  ]
}
```

**Fields that differ by brand.** GTIN barcodes, promotion windows, dimensions, weight, origin, `traceability`, `certified_materials`, `care`, `on_special`, `is_buyable`, `back_soon`, `visibility`, `availability_date`, `join_life`, the sustainability node, `categories`, `additional_info`, `keywords`, `assembly_url` and every `*_en` name column come from the six shared-catalogue brands, not from Zara. `first_visible_date`, `grid_position`, `is_pinned`, `seo_keyword`, `section_name`, `category_id`, `category_name`, `size_guide`, `available_color_names` and per-SKU `demand` come from Zara. Massimo Dutti publishes no GTIN barcodes at all and very few descriptions, which is why its own single-brand Actor is priced lower. `traceability` and `certified_materials` are present in the schema but frequently empty — treat them as a bonus, not a guarantee.

### 💼 Use Cases & Examples

#### 1. Cross-brand price architecture

**A pricing team wants to see how the group ladders its brands.**

**Input:** all seven brands, one market, full sweep
**Output:** price, family and subfamily taxonomy, `brand` on every row
**Use:** average price per family per brand, where the brands overlap and where they do not

#### 2. Group-wide markdown tracking

**A retail analyst wants the whole group's discount calendar, not one brand's.**

**Input:** a weekly run across all seven brands
**Output:** `old_price`, `discount_pct`, `on_special`, and per-SKU promotion start and end dates
**Use:** when each brand starts and ends a markdown, and how deep it goes

#### 3. Barcode-keyed catalogue matching

**A marketplace needs to match Inditex products against its own inventory.**

**Input:** the six shared-catalogue brands, one market
**Output:** GTIN barcodes per SKU, with price, size and stock
**Use:** joining to any other retail dataset without fuzzy name matching

#### 4. Assortment overlap analysis

**A category manager wants to know where two sibling brands compete with each other.**

**Input:** two brands, same market, same categories
**Output:** one table with identical taxonomy columns for both
**Use:** range overlap by family, price gaps at the same product type

#### 5. Multi-market expansion research

**A team is choosing which market to enter and wants the group's local pricing.**

**Input:** the same brands run across several `country` values
**Output:** each market's own prices and currency, with `country` on every row
**Use:** local price levels, assortment differences, currency-adjusted benchmarks

#### 6. Sustainability and composition reporting

**A researcher wants fibre composition across the group's catalogue.**

**Input:** all seven brands, full sweep
**Output:** `composition` per garment part on every brand; `care`, `join_life` and the sustainability node on the six shared-catalogue brands
**Use:** fibre mix by brand and family, synthetic share, care-label analysis

### 🔗 Integration Examples

#### JavaScript/Node.js

```javascript
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });

const run = await client.actor('sian.agency/inditex-product-scraper').call({
  brands: ['zara', 'bershka', 'stradivarius'],
  mode: 'overview',
  country: 'ES',
  allCategories: true,
  maxItems: 600,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Got ${items.length} products across ${new Set(items.map((i) => i.brand)).size} brands`);
```

#### Python

```python
from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')

run = client.actor('sian.agency/inditex-product-scraper').call(
    run_input={
        'brands': ['zarahome', 'oysho'],
        'mode': 'overview',
        'country': 'FR',
        'allCategories': True,
        'maxItems': 400,
    }
)

for item in client.dataset(run['defaultDatasetId']).iterate_items():
    print(item['brand'], item['name'], item['price'], item['currency'])
```

#### cURL

```bash
curl -X POST 'https://api.apify.com/v2/acts/sian.agency~inditex-product-scraper/runs?token=YOUR_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
  "brands": ["pullandbear"],
  "mode": "overview",
  "country": "GB",
  "allCategories": true,
  "maxItems": 200
}'
```

#### Automation Tool Workflows (n8n, Zapier, Make, etc.)

1. **Trigger**: schedule or webhook
2. **HTTP Request**: start a run and wait for the dataset
3. **Process**: group by `brand` and `family_name_en`, diff against last week
4. **Action**: write to your warehouse, or alert on new markdowns across the group

### 📊 Performance & Pricing

[💰 View current pricing](https://apify.com/sian.agency/inditex-product-scraper?fpr=sian)

#### Performance

- \~20 products per second on the category sweep
- Brands run sequentially into one dataset; **Max Products** is a run total across all of them
- 512 MB memory, no proxy, no browser — direct API reads
- Duplicates removed before anything is charged, so you pay once per distinct product

#### How you are billed

Three events, and the `source` column on every row tells you which one applied. Your invoice reconciles against the dataset, brand by brand.

| What you ran | `source` on the row | Event | Price |
|---|---|---|--:|
| Run start | — | Actor start | $0.005 once |
| Category sweep | `overview` | Scraped product | **$0.0035** |
| Pasted URL, or a sitemap row | `detail` | Scraped product detail | $0.0105 |
| Zara sweep row plus its SKU fetch (`withSizes`) | `overview+detail` | both | $0.014 |

**Two switches move a row onto the higher-priced event.** **Also Crawl the Sitemap** cannot use the bulk endpoint — it is one request per product, so those rows bill at $0.0105 instead of $0.0035. **Fetch Sizes & SKUs** on Zara genuinely costs both calls and charges both, so those rows cost $0.014. The category sweep is the default and the cheap path; use the sitemap when completeness matters more than price.

#### Cost examples

- **1,000 products across three brands, category sweep**: $3.505
- **7,000 products, all seven brands, category sweep**: $24.505
- **500 pasted product URLs**: $5.255

Prices shown are the BRONZE tier. They step down at SILVER, GOLD, PLATINUM and DIAMOND, so heavier use costs less per row.

#### What it replaces

The only vendor covering all seven brands sells them as seven monthly rentals: $24.99 for Zara plus six at $9.99, **$84.93 a month** before you have scraped anything. Here, 7,000 products across all seven brands costs $24.51 and you pay nothing in the months you do not run it.

**If you only want one brand, use its own Actor.** The single-brand Actors run $0.0025–$0.003 per product, so per row they are cheaper than this one. What the extra buys here is the normalization: seven catalogues land in one dataset with one schema and a `brand` column, from a single run you schedule once. That is worth paying for when you need to compare the brands against each other. If you do not, run the single-brand Actor instead.

### ❓ Frequently Asked Questions

**Q: Do all seven brands really return the same columns?**
A: Yes — 58 columns on every row, whichever brand produced it. Columns only one catalogue carries come back null on the others, and the table above says which is which.

**Q: Is Max Products per brand or per run?**
A: Per run. Selecting seven brands with a cap of 100 gives you 100 products in total, not 700. The run splits that total between the brands rather than letting the first one swallow it, so all seven are represented; a brand with less stock than its share passes the remainder on. Set it to `0` for no cap.

**Q: Can I mix brands in `detail` mode?**
A: Yes, if you paste full product URLs — each is routed to the brand its hostname names. A bare numeric id goes to the first brand in your list, so do not mix bare ids across brands.

**Q: Why does Zara come back without barcodes?**
A: Zara's catalogue does not publish GTIN codes; the six shared-catalogue brands do. Massimo Dutti is the one exception among those six — it publishes no barcodes and very few descriptions.

**Q: Can I search by keyword?**
A: No, and neither can anything else. None of these sites exposes a server-side search endpoint. Name a category or sweep them all and filter the results yourself.

**Q: Does it return customer reviews?**
A: No. No Inditex brand publishes ratings or reviews through these catalogues, so no scraper can return them.

**Q: Which markets are supported?**
A: Not the same number for every brand. The six shared-catalogue brands are open in ~216 markets each; zara.com is open in 96. The Market dropdown lists the union of 220, so a market outside Zara's 96 works for the other six but stops a run that includes Zara — deselect Zara, or pick a market Zara serves. Prices come back in that market's own currency with the correct divisor applied.

**Q: What output formats are available?**
A: JSON, CSV and Excel, exported straight from the dataset. There is also a run summary in the key-value store with the counts for the run.

### 🐛 Troubleshooting

**One brand returned nothing while the others worked**

- That brand's category may be a container rather than a product grid — those return no ids
- The run continues and warns rather than failing, so check the log for which brand was skipped

**The run summary says `complete: false`**

- Some fetches failed for transport reasons, so the dataset is short. Re-run to fill the gap
- Ids reported as gone are a different thing: those products no longer exist in that market

**Sort stopped the run**

- Zara and the six shared-catalogue brands use different sort ids. `default` is the only option that works for every brand at once

**The run stopped on the market, not the data**

- The Market dropdown is the union across the seven brands. zara.com is open in 96 markets, the other six in ~216, so a market outside Zara's 96 stops a run that includes Zara
- Deselect Zara in **Brands**, or pick a market Zara serves

**A bare product id went to the wrong brand**

- Bare ids go to the first brand in your list. Paste full URLs when mixing brands

**Sitemap mode is slower and dearer than expected**

- That is the trade: it is one request per product instead of one per hundred, and those rows bill at the detail rate. Use the category sweep unless you need the products no category lists

***

### ⚖️ Is it legal to scrape data?

Our actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our actors, when used for ethical purposes by Apify users, are safe.

However, you should be aware that your results could contain personal data. Personal data is protected by the **GDPR** in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers.

You can also read Apify's blog post on the [legality of web scraping](https://blog.apify.com/is-web-scraping-legal/).

**Trademarks.** Zara, Zara Home, Bershka, Massimo Dutti, Stradivarius, Oysho and Pull\&Bear are trademarks of Industria de Diseño Textil, S.A. This Actor is an independent tool. It is not affiliated with, endorsed by or sponsored by Inditex or any of its brands, and it reads only publicly available catalogue pages.

***

### 🤝 Support

[![Telegram Support](https://img.shields.io/badge/Telegram-Support%20Group-0088cc?logo=telegram)](https://t.me/+vyh1sRE08sAxMGRi)

**Join our active support community**

- 🐛 Found a bug? File an issue in the Apify Console Issues tab
- ⭐ Loving the tool? Leave a 5-star review — it helps us build more
- Check [SIÁN Agency Store](https://apify.com/sian.agency?fpr=sian) for more automation tools
- 📧 <apify@sian-agency.online>

***

**Built by [SIÁN Agency](https://www.sian-agency.online)** | **[More Tools](https://apify.com/sian.agency?fpr=sian)**

# Actor input Schema

## `brands` (type: `array`):

Which of the seven Inditex brands to scrape. Brands run one after another into the same dataset and every row carries a `brand` column, so a combined dataset stays sortable.

The item cap below is a **run total** across every brand you select, not a per-brand allowance. The run splits it evenly between the brands still to go, so a small cap still returns rows for each of them, and a brand with less stock than its share hands the remainder to the brands after it.

## `mode` (type: `string`):

**Overview** walks the category tree (and optionally the sitemap) and returns every product it finds.

**Detail** takes the product URLs you paste below and returns one full record each. Detail mode needs at least one URL or it stops immediately.

## `country` (type: `string`):

Two-letter store country, e.g. `DE`, `ES`, `FR`, `GB`, `JP`.

This list is the **union across the seven brands**, not a list every brand serves. The six shared-catalogue brands are open in ~216 markets each; **zara.com is open in only 96**.

Pick a market outside Zara's 96 with Zara selected and the run stops. Either drop Zara from the brand list or choose a market it serves. `GB` is safe — zara.com serves it as `uk`.

## `language` (type: `string`):

Two-letter language for names and descriptions. Leave empty and each store serves its own default for the market you picked.

An unsupported code stops the run with the list the market does support.

## `allCategories` (type: `boolean`):

Walk every category in the tree. On by default so a bare run returns data.

**This wins over the single category below.** Turn it off to scrape one category.

Categories that are containers rather than product grids return nothing and are skipped, which is normal — the run summary counts them under `categoriesSkipped`.

## `category` (type: `string`):

One category, by slug or numeric id. The slug is the last path segment of a category URL on the brand site.

Turn off **Sweep All Categories** above or this field is ignored.

## `fromSitemap` (type: `boolean`):

After the categories, read the product sitemap and fetch everything the categories missed.

This is the expensive half of a full-catalogue run: sitemap products are fetched one at a time rather than in bulk, so those rows are charged as the "Scraped product detail" event at $0.0105 instead of $0.0035 — 3× the category-sweep price.

Use it when you want the complete catalogue rather than what merchandising currently surfaces.

## `maxItems` (type: `integer`):

Stop after this many products. `0` means no cap — the run goes until the catalogue is exhausted.

Duplicates are removed before the cap is counted, so the number you set is the number of distinct products you get.

## `withSizes` (type: `boolean`):

Zara's grid ships no sizes and no description. Switch this on and the actor fetches the detail record for every product, ten ids per request, and merges it into the grid row.

That adds per-SKU sizes, stock, the description, Zara's demand signal, twinned SKUs and the first-seen date.

Applies to Zara only. The other six brands already carry sizes on the overview call.

Billing: these rows genuinely cost both calls and are charged both events — $0.0035 + $0.0105 = $0.014 per product.

## `productUrls` (type: `array`):

One product per line — a full product URL, or a bare product id.

URLs are routed to the brand their hostname names. A bare id goes to the first brand in your list, so mix brands only when you paste full URLs.

## `sort` (type: `string`):

Zara and the six shared-catalogue brands use different sort ids, so pick one that matches the brands you selected. `Catalogue order` is the only option that works for every brand at once.

Both sites publish pre-sorted id lists rather than accepting a sort parameter. A sort id the selected brand does not use stops the run and prints the valid ones.

On Zara a category that lacks the chosen sort stops the run; on the other six it simply returns catalogue order.

## `filters` (type: `array`):

Facet filters as `GROUP=VALUE`, one per line. Values in the same group are OR-ed, different groups are AND-ed.

On the six shared-catalogue brands use the shortcuts `color=`, `size=`, `category=` and `discount=`, or the raw facet group id. Zara uses its own ids (`color`, `size`), which this field passes through unchanged.

Facet groups are per category: one the category does not offer stops the run and prints the ones it does.

## `minPrice` (type: `integer`):

Drop products cheaper than this, in major units of the market currency. Leave empty for no floor.

## `maxPrice` (type: `integer`):

Drop products dearer than this, in major units of the market currency. Leave empty for no ceiling.

## Actor input object example

```json
{
  "brands": [
    "zara",
    "zarahome",
    "bershka",
    "massimodutti",
    "stradivarius",
    "oysho",
    "pullandbear"
  ],
  "mode": "overview",
  "country": "DE",
  "language": "",
  "allCategories": true,
  "fromSitemap": false,
  "maxItems": 100,
  "withSizes": false,
  "sort": "default",
  "filters": [
    "color=Gelb",
    "size=M"
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

One row per product: price, old price, currency, discount, images, taxonomy, colours, per-SKU sizes and stock.

## `runSummary` (type: `string`):

Your run at a glance: products returned, ids that no longer exist, fetches that failed and how to retry them, how each row was fetched, and an itemized statement of what you paid.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "brands": [
        "zara",
        "zarahome",
        "bershka",
        "massimodutti",
        "stradivarius",
        "oysho",
        "pullandbear"
    ],
    "mode": "overview",
    "country": "DE",
    "language": "",
    "maxItems": 100,
    "sort": "default"
};

// Run the Actor and wait for it to finish
const run = await client.actor("sian.agency/inditex-product-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "brands": [
        "zara",
        "zarahome",
        "bershka",
        "massimodutti",
        "stradivarius",
        "oysho",
        "pullandbear",
    ],
    "mode": "overview",
    "country": "DE",
    "language": "",
    "maxItems": 100,
    "sort": "default",
}

# Run the Actor and wait for it to finish
run = client.actor("sian.agency/inditex-product-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "brands": [
    "zara",
    "zarahome",
    "bershka",
    "massimodutti",
    "stradivarius",
    "oysho",
    "pullandbear"
  ],
  "mode": "overview",
  "country": "DE",
  "language": "",
  "maxItems": 100,
  "sort": "default"
}' |
apify call sian.agency/inditex-product-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sian.agency/inditex-product-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9O7VNDAoCLh9Pti6x/builds/JzvXWzjoAeDmFfgfV/openapi.json
