# Hepsiburada Products Scraper (`scrapyx/hepsiburada-products-scraper`) Actor

Product listings from Hepsiburada, Turkey's second-largest marketplace: price in lira, instalment offer, rating, reviews and badges. The price is read from its own node -- the first number on a card is the instalment, and a naive read reports a 32,999 TL laptop at 5,499.

- **URL**: https://apify.com/scrapyx/hepsiburada-products-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.05 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hepsiburada Products Scraper

Product listings from **Hepsiburada**, Turkey's second-largest marketplace:
price in lira, the instalment offer, rating and review count, badges, image and
product URL — for any search term.

HTTP only, no browser, no key. Off Hepsiburada's own server-rendered listing.

### Input

| field | what it does |
| --- | --- |
| `searchTerms` | `laptop`, `kulaklık`, `iphone 15` — one target per term. |
| `maxItems` | Overall cap. One term tops out at ~1,800 (see below). |
| `maxConcurrency`, `minRequestInterval` | Pacing. Pages are 3.5 MB each. |

### Four things about this data worth knowing

#### 1. The first price on the card is the instalment, not the price

A card's text, in the order it appears, for a 32,999 TL laptop:

```
'Lenovo IdeaPad Slim 3 …', 'Ürün puanı 4.9, 23 değerlendirme', '4,9',
'Kampanya: Peşin fiyatına 6 x 5.499 TL', '6 x', '5.499 TL',   ← first "N TL"
'32.999', 'TL'                                                 ← the price
```

"Peşin fiyatına 6 x 5.499 TL" is *six instalments at the cash price*. A
scraper that takes the first string containing "TL" reports **5,499 TL** for a
**32,999 TL** laptop — wrong by about 6×, and nothing about 5,499 looks wrong
for a laptop. This Actor reads the price from its own node only, and keeps the
instalment beside it as `installmentCount` × `installmentAmount`.

#### 2. Two id spaces on one page

```
…-83k700pstr-p-HBCV0000FC0KHB      a SKU   (one specific variant)
…-bilgisayar-pm-HBC0000D9H7T8      a MODEL (a product with variants)
```

Half the cards on a page are one, half the other. A parser keyed on `-p-`
alone returns null ids for the model half, and a run that de-duplicates on
that id then collapses them all into one row. Both are read; every row carries
`idType`, and the summary counts `skuRows` and `modelRows` so a regression in
either is visible.

#### 3. Turkish numbers, with the decimals in a separate node

`68.238` + `,40` + `TL` is **68,238.40** lira: `.` is the thousands separator,
`,` the decimal mark, and the `,40` lives in a nested node whose text is just
`TL` when there are no decimals. `float("68.238")` is wrong by 1000× and still
looks like a plausible price for something. Every row keeps `priceRaw` so the
arithmetic is auditable.

#### 4. Page 50 answers HTTP 403 — and it is not a block

Page 49 returns 36 cards. Page 50 returns **403** with a 1,753-byte body; so do
pages 500 and 9999, while page 1 keeps answering 200 in the same session. It is
Hepsiburada's way of saying "no such page", but it reads exactly like a WAF
kicking in mid-crawl. The wall is pinned at 49 and never requested — roughly
1,800 products per term — and reaching it is reported as `ceilingHit`.

### Output

- **`PRODUCT`** — `productId` + `idType`, `title`, `price` + `priceRaw` +
  `currency`, `installmentCount` + `installmentAmount` + `campaignText`,
  `rating`, `reviewCount`, `badges`, `imageUrl`, `productUrl`.
- **`SEARCH_SUMMARY`** — `productsReturned`, `matchCountAvailable: false`,
  `skuRows`, `modelRows`, `rowsWithoutId`, `duplicateRowsDropped`,
  `ceilingHit`, `stoppedReason`.
- **`ERROR`** — one row naming what went wrong, instead of a silent empty.

### Technical notes

- **Every class name carries a build hash** (`price-module_finalPrice__LtjvY`)
  that changes on deploy. Every selector matches on the stable stem with
  `[class*=…]`, never on a full class name.
- **One search in a handful comes back as HTTP 206** with the complete page. A
  200-only classifier files that as fatal and loses the term; 2xx is treated as
  success here and the card count decides.
- **Two kinds of empty page.** A nonsense term returns zero cards *with* a
  "bulunamadı" marker — an honest no-results. An empty term returns zero cards
  *without* it — the search page with nothing searched. The empty term is
  refused locally; zero cards without the marker is reported as
  `page_shape_changed` rather than as "no results".
- **No keyword-match check**, deliberately. Titles are Turkish
  ("Taşınabilir Bilgisayar"), queries often English ("laptop"), and the failure
  such a check guards against elsewhere — a dropped keyword answered with an
  unrelated catalogue — was measured not to exist here. A check with no
  true-positive case and a routine false-positive case is worse than none.
- **robots.txt** checked 2026-09-16: no AI-bot group, no blanket disallow,
  `/ara?q=` allowed. The platform inventory was checked before a line was
  written.

### Known limits

- The listing carries **no struck-through price**, **no merchant name** and
  **no result total** — checked on a discount-sorted search too. Product detail
  pages are not read.
- \~1,800 products per search term. Slice with narrower terms to go wider.
- Prices in lira only.

# Actor input Schema

## `searchTerms` (type: `array`):

What to search for on Hepsiburada - 'laptop', 'kulaklık', 'iphone 15'. Each term becomes its own target with its own summary row. An empty term is refused here: Hepsiburada answers it with the search page and zero products, which is indistinguishable from a broken parser.

## `maxItems` (type: `integer`):

Overall cap on product rows across every search term. Hepsiburada serves 36-37 products per page and answers page 50 with HTTP 403 - that is the end of the listing, not a block - so a single term tops out at about 1,800 products. A run that reaches the wall reports ceilingHit.

## `maxConcurrency` (type: `integer`):

Parallel requests across search terms. Pages are 3.5 MB each, so keep this modest.

## `minRequestInterval` (type: `integer`):

Politeness delay between request starts. It paces starts only and does not hold a concurrency slot, so raising it slows the run without idling workers.

## `proxyConfiguration` (type: `object`):

Optional. Hepsiburada answered 200 cold from a plain IP. If a cloud run reports fetch\_failed on page 1, enable Apify's free datacenter proxy first - some sites rate-limit Apify's own egress address and nothing else.

## Actor input object example

```json
{
  "searchTerms": [
    "laptop"
  ],
  "maxItems": 200,
  "maxConcurrency": 2,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "laptop"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/hepsiburada-products-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchTerms": ["laptop"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/hepsiburada-products-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "laptop"
  ]
}' |
apify call scrapyx/hepsiburada-products-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/hepsiburada-products-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/a9e5NxqJF4tfl9mGM/builds/ZO9WmFdwF8R7akNPf/openapi.json
