# Jarir Books & Electronics Catalog Scraper (`coolinbex/jarir-books-electronics-scraper`) Actor

Extract public Jarir Saudi books and electronics with ISBNs, authors, SKUs, brands, specifications, prices, availability, images, and category data.

- **URL**: https://apify.com/coolinbex/jarir-books-electronics-scraper.md
- **Developed by:** [coolinbex](https://apify.com/coolinbex) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Jarir Books & Electronics Catalog Scraper

Extract public Jarir Saudi Arabia catalog pages for market research, inventory discovery, price monitoring, and content enrichment. The Actor uses only browser-visible public pages; it does not log in, call private APIs, defeat access controls, or bypass CAPTCHA pages.

### What it extracts

Each dataset record can include the product title, Jarir URL and canonical URL, category, SKU/product ID, ISBN, authors, brand, current and old price, currency, availability, description, public images, and visible specification-like label/value pairs. Books are especially useful for ISBN/author data; electronics records are useful for SKU/brand/specification/price comparisons.

### Input

`startUrls` is a request list of public `www.jarir.com` category, search, or product URLs. The default uses the public electronics category because it currently exposes a product record reliably; it is intentionally limited to one page and one product without detail-page enrichment so Apify health tests complete quickly. Set `scrapeDetails: true` for richer ISBN, specification, and availability fields. `category` can be `all`, `books`, or `electronics`. `maxItems` (1–1000), `maxPages` (1–20), `scrapeDetails`, `includeImages`, `includeSpecifications`, `maxConcurrency` (1–2), and optional authorized `proxyConfiguration` control bounded usage.

Example:

```json
{
  "startUrls": [{ "url": "https://www.jarir.com/sa-en/electronics.html" }],
  "category": "electronics",
  "maxItems": 50,
  "maxPages": 2,
  "scrapeDetails": true
}
```

### Output and limits

Records are written incrementally to the default dataset. `status` is `ok` or `partial`; `errorMessage` is populated for a failed detail request. A `SUMMARY` key-value record reports pages, unique products, retries/failures, blocked pages, partial records, and source-change signals. The Actor stops at the configured item/page/request bounds, uses one-to-two browser sessions, 1.5 seconds same-domain pacing, 45-second navigation timeouts, and two retries per request. Empty or blocked runs are reported explicitly with a summary and do not fabricate products.

Jarir’s storefront is client-rendered and selectors/data shapes can change. A non-`ok` record or `sourceChanges` summary signal should be reviewed before automated downstream use. Pagination and product discovery depend on public links rendered on the supplied pages.

### Runtime and cost

Browser rendering is the main cost driver. A small 10–50 product run is generally the economical starting point; detail pages roughly double page requests. Use `scrapeDetails: false`, low `maxItems`, and one concurrency for previews. Apify platform compute, proxy traffic (if selected), storage, and any applicable plan minimums are billed by Apify; this Actor does not add a separate fee. For a commercial price, a practical model is per successful product record, with a higher tier for detail enrichment and an explicit pass-through for proxy/compute costs.

### Safe and compliant use

Use only public pages you are authorized to collect, honor Jarir’s terms and robots guidance, keep concurrency and page bounds low, and do not use this Actor to access accounts, personal data, checkout flows, or restricted endpoints. Validate price/availability before business decisions because storefront values change.

# Actor input Schema

## `startUrls` (type: `array`):

Public www.jarir.com Saudi Arabia category, search, or product URLs.

## `maxItems` (type: `integer`):

Maximum unique product records to write.

## `maxPages` (type: `integer`):

Bound pagination depth for each category or search URL.

## `scrapeDetails` (type: `boolean`):

Open each discovered public product page for richer fields. Disabled by default for fast health checks.

## `category` (type: `string`):

Keep only books, electronics, or both.

## `includeImages` (type: `boolean`):

Keep public product image URLs when available.

## `includeSpecifications` (type: `boolean`):

Keep visible product specification label/value pairs when available.

## `maxConcurrency` (type: `integer`):

Keep low for responsible public-page crawling.

## `proxyConfiguration` (type: `object`):

Optional authorized Apify proxy configuration for larger runs.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.jarir.com/sa-en/electronics.html"
    }
  ],
  "maxItems": 1,
  "maxPages": 1,
  "scrapeDetails": false,
  "category": "all",
  "includeImages": true,
  "includeSpecifications": true,
  "maxConcurrency": 1
}
```

# Actor output Schema

## `products` (type: `string`):

Structured Jarir product records.

## `summary` (type: `string`):

Counts, status, and source-change signals for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("coolinbex/jarir-books-electronics-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("coolinbex/jarir-books-electronics-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call coolinbex/jarir-books-electronics-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,coolinbex/jarir-books-electronics-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/FBgRtxbood2yM4m9Y/builds/z1OVVro6KKKehI7eH/openapi.json
