# Hamiz Multi-Scraper (`hamza325/hamiz-multi-scraper`) Actor

Scrape hamiz.com, Algeria's buy-and-sell marketplace, with lightweight HTTP requests against the site's own search, category, and seller pages. No login and no browser automation required.

- **URL**: https://apify.com/hamza325/hamiz-multi-scraper.md
- **Developed by:** [Hamza Abbad](https://apify.com/hamza325) (community)
- **Stats:** 2 total users, 1 monthly users, 91.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hamiz Multi-Scraper

Scrape [hamiz.com](https://hamiz.com) — Algeria's buy-and-sell marketplace — with lightweight HTTP requests against the site's own search, category, and seller pages. No login and no browser automation required. Every record is a flat, normalized object ready to feed straight into a spreadsheet, a database, or another Actor downstream.

Hamiz is a fixed-price marketplace: sellers list products with a set price in Algerian Dinar, a stock state, and delivery handled by the seller. There is no bargaining concept and no negotiable-price flag — the price you see is the price. The site is trilingual (Arabic, French, English); the **Language** (`locale`) input picks which storefront language your run uses, which also decides the language prefix of the result URLs.

### What you can scrape

The Actor drives five modes from a single **What to scrape** (`startMode`) input. Pick the one that matches what you want.

| Key (`startMode` value) — Console label | What it returns | Best for |
| --- | --- | --- |
| `search` — **Search by keyword** | Every product matching a free-text query, walking the site's own search result pages (20 per page; card rows: id, title, SKU, price, image) | Market research, price tracking, product discovery |
| `category` — **Browse a category** | All products in one category, paginated at 20 per page | Browsing a niche (car accessories, phones, home goods…) |
| `directUrls` — **Specific product URLs** | A specific list of product pages, fully enriched: description, SKU, special and regular prices, discount percent, stock state, seller card, image gallery, rating, delivery note | Following up on alerts, checking one item's details, resuming a previous run |
| `seller` — **Seller listings** | Every product a seller has listed, one row per product, paginated | Tracking a shop's inventory, monitoring a seller |
| `categories` — **Category tree snapshot** | Every category node from the site menu (hundreds of nodes) with its depth and parent chain | Building a taxonomy, seeding category runs, site mapping |

> All modes share the same flat row shape; each mode populates only the fields it has data for (the others are `null` or empty). Grid modes (`search` / `category` / `seller`) fill the card fields; `directUrls` fills the full projection; `categories` emits category nodes marked with `__type: "hamiz_category"`. Every row is one product (or one category node). Use the key names (`search`, `category`, `directUrls`, `seller`, `categories`) in JSON/API calls — the **What to scrape** (`startMode`) picker shows the Console labels.

### Why this Actor

- **Search that works without a browser.** Keyword search walks the same result pages a visitor sees. The search page is the site's most protected surface, so every request replays a real navigation: a warmed session with the site's cookies, the same-origin referer a page-to-page click would send, and locale headers. No browser automation, no CAPTCHA solving, fast on cold start and cheap to run at scale.
- **Flat, uniform output.** Every row is a flat object with no nested structures — the dataset table shows clean values instead of "N fields" badges. Specification tables become two parallel lists (`specNames` / `specValues`) whose entries pair up by position.
- **Real discount math.** Products on sale carry both prices plus the percent: `price` (what you pay), `regularPrice` (before the sale), and `discountPct` (e.g. `28` for "-28%"). `currency` is always `"DZD"`.
- **Seller context where it exists.** Product pages name the seller ("Sold by …"), show a verified-seller mark, and link the storefront — the Actor captures all three, plus the seller avatar. **Seller listings** (`seller`) then walks that seller's whole inventory.
- **You only ever paste URLs.** Every reference field takes a full page URL, straight from the browser's address bar. No IDs, no slugs, no values copied out of the address bar — and the Actor tells you exactly what it expected if a URL points at the wrong kind of page.

### Quick start

1. Open this Actor in Apify Console.

2. Click **Try actor** and set **What to scrape** (`startMode`) to **Search by keyword** (`search`), plus a **Search keyword** (`searchQuery`) — the minimum valid JSON is:

   ```json
   { "startMode": "search", "searchQuery": "iphone" }
   ```

3. Click **Start**. Results land in the **Dataset** tab, one JSON object per product.

The smallest valid input runs a keyword search across the whole site, with sensible defaults for everything else.

### Tutorial: real workflows

#### 1. Track prices for a keyword

Search returns the site's own card data — id, title, SKU, final price, image — 20 products per page, walking every result page until the result count on the first page is covered:

```json
{
  "startMode": "search",
  "searchQuery": "climatiseur",
  "locale": "fr",
  "maxItems": 500,
  "maxPages": 20
}
```

Schedule this daily and diff the `price` column per `id` to catch price moves.

#### 2. Enrich the interesting hits

Search rows are deliberately lean (no description, no seller, no stock). Take the `url` values you care about and run them back through the Actor for the full projection:

```json
{
  "startMode": "directUrls",
  "startUrls": ["https://hamiz.com/en/4-piece-car-window-sunshade-set-black-s6223.html"]
}
```

This fills `description`, `sku`, `regularPrice`, `discountPct`, `availability`, the seller card (`sellerName`, `sellerVerified`, `sellerProfileUrl`, `sellerAvatar`), the full `images` gallery, `rating`, `questionCount`, `deliveryInfo`, and the spec table (`specNames` / `specValues`).

#### 3. Walk a seller's inventory

Paste any seller storefront URL to get everything they list:

```json
{
  "startMode": "seller",
  "sellerUrl": "https://hamiz.com/en/dima-top",
  "maxItems": 500
}
```

#### 4. Map the taxonomy

One run with **Category tree snapshot** (`categories`) emits every category node with its `depth` and `parentPath` (e.g. `Categories › Automobile › …`). Use the node URLs as **Category page URL** (`categoryUrl`) inputs for follow-up browse runs.

### Limits

- **Rate limiting.** The site sits behind bot protection that watches request patterns. The default pace (30 requests per minute) is deliberately polite — raising it makes blocks more likely, not runs faster.
- **Search recall.** Keyword search walks the site's own result pages, so the row count matches the result count the site shows for the query. For exhaustive coverage of a niche, combine a search run with a category browse.
- **No contact details.** Hamiz exposes no phone numbers, WhatsApp contacts, or emails on product or seller pages — buyer questions go through the site's own Q\&A tab (captured as `questionCount`). There is no contact-info toggle because there is nothing to reveal.
- **Seller-entered text is never translated.** Titles and descriptions stay in whatever language the seller wrote (often Arabic or French) regardless of the **Language** (`locale`) input — that input only affects interface labels and URL prefixes.
- **Live counters are point-in-time.** Ratings and question counts reflect the moment of the run.

### FAQ

**Do I need a proxy?**
No. On the Apify platform the Actor automatically probes proxy addresses and locks onto one that returns data, falling back to direct egress when none is needed.

**Why do some fields come back `null`?**
Each mode fills only what its surface provides. Descriptions, stock states, and seller cards need a product page — run those URLs through **Specific product URLs** (`directUrls`) to fill them.

**What does `discountPct: null` mean?**
The product has no special offer — `price` is the only price. It does not mean "no discount data".

**How do I pair `specNames` with `specValues`?**
By position: entry N of `specNames` (e.g. `"Color"`) describes entry N of `specValues` (e.g. `"Black"`).

**The run stopped before `maxItems` — is that an error?**
No. Grids end when the site runs out of products (an empty page stops pagination), and the run ends when every queued page is done. Re-running with the same settings is idempotent thanks to built-in dedupe.

# Changelog

This Actor's version history is a separate document: https://apify.com/hamza325/hamiz-multi-scraper/changelog.md

# Actor input Schema

## `startMode` (type: `string`):

How to drive the scraper. **Search by keyword** (`search`) walks the site's own search result pages (card rows: id, title, price, image, 20 per page). **Browse a category** (`category`) and **Seller listings** (`seller`) walk paginated product grids. **Specific product URLs** (`directUrls`) enriches a hand-picked list of product URLs with the full detail data (description, SKU, stock state, seller card, images, rating, delivery note). **Category tree snapshot** (`categories`) emits every category node from the site menu with its depth and parent chain.

## `searchQuery` (type: `string`):

Used when **What to scrape** (`startMode`) = **Search by keyword** (`search`). Free-text query, exactly as the site's own search box accepts it (e.g. `iphone`, `climatiseur`, `abaya`).

## `categoryUrl` (type: `string`):

Used when **What to scrape** (`startMode`) = **Browse a category** (`category`). Paste the category page URL straight from the browser's address bar (e.g. `https://hamiz.com/en/automobile.html`). Nested category paths work the same way.

## `sellerUrl` (type: `string`):

Used when **What to scrape** (`startMode`) = **Seller listings** (`seller`). Paste the seller storefront URL straight from the browser's address bar (e.g. `https://hamiz.com/en/dima-top`). The run returns every product the seller has listed, one row per product.

## `startUrls` (type: `array`):

Used when **What to scrape** (`startMode`) = **Specific product URLs** (`directUrls`). The product page URLs, pasted straight from the browser's address bar (e.g. `https://hamiz.com/en/4-piece-car-window-sunshade-set-black-s6223.html`). Each URL is scraped with the full data the site offers: description, SKU, special and regular prices, stock state, seller card, image gallery, rating, and delivery note.

## `maxItems` (type: `integer`):

Stop after pushing this many rows. Applies in every mode. 0 = no cap.

## `maxPages` (type: `integer`):

Stop after this many page requests. One page = one request (search and grid pages carry 20 products each). Use 0 for no cap. The scraper stops on its own when a grid runs out.

## `maxRequestsPerMinute` (type: `integer`):

Rate limit per worker. Hamiz sits behind bot protection that watches request patterns — stay at or below the default so the Actor stays viable long-term.

## `locale` (type: `string`):

Request language. Sets the site language prefix (`/en/`, `/fr/`, `/ar/`) — result URLs and interface labels follow it. Product titles and descriptions are seller-entered and never translated.

## `dedupe` (type: `boolean`):

Skip products whose ID has already been pushed in this run. Duplicates are dropped at the request-queue level too, so re-running the Actor with the same settings is idempotent.

## `debugIncludeRaw` (type: `boolean`):

If true, each pushed row includes a `raw` field with the unprocessed data extracted from the page (the raw HTML is never included).

## Actor input object example

```json
{
  "startMode": "search",
  "searchQuery": "iphone",
  "categoryUrl": "https://hamiz.com/en/automobile.html",
  "sellerUrl": "https://hamiz.com/en/dima-top",
  "startUrls": [],
  "maxItems": 200,
  "maxPages": 20,
  "maxRequestsPerMinute": 30,
  "locale": "fr",
  "dedupe": true,
  "debugIncludeRaw": false
}
```

# Actor output Schema

## `products` (type: `string`):

Default dataset containing every product the Actor pushed during the run — one flat row per product, in the Products view.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "iphone",
    "startUrls": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("hamza325/hamiz-multi-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "iphone",
    "startUrls": [],
}

# Run the Actor and wait for it to finish
run = client.actor("hamza325/hamiz-multi-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "iphone",
  "startUrls": []
}' |
apify call hamza325/hamiz-multi-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hamza325/hamiz-multi-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KXJ7NgyoWHTfEgevh/builds/Ax7huKgZHzQ56JBMv/openapi.json
