# Dubizzle UAE Vehicles Scraper (`husseinelabeedy/dubizzle-uae-vehicles-scraper`) Actor

Scrape used car listings from Dubizzle UAE quickly using its Algolia search index. Collect vehicle details, prices, locations, sellers, specifications, images, and listing URLs without opening individual listing pages. Supports year ranges, concurrent scraping, and direct Apify Dataset output.

- **URL**: https://apify.com/husseinelabeedy/dubizzle-uae-vehicles-scraper.md
- **Developed by:** [AL-Hussein EL-Abeedy](https://apify.com/husseinelabeedy) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 vehicles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dubizzle UAE Cars Scraper

Scrape used car listings from **Dubizzle UAE** quickly and efficiently using Dubizzle's Algolia search index.

The Actor collects vehicle listings directly from Dubizzle's search infrastructure without opening individual car listing pages. This makes it suitable for collecting large datasets of used cars across different vehicle years.

### Performance

The scraper can collect the full available set of **34,000+ car listings in approximately 8+ minutes** in a full-range run.

Performance depends on the selected year range, number of available listings, network conditions, and configured concurrency.

**Example:**

- Listings collected: **34,000+ cars**
- Runtime: **8+ minutes**
- Requests: Concurrent
- Output: Apify Dataset
- Duplicate listings: Automatically removed

### Features

- Scrapes used car listings from Dubizzle UAE
- Covers multiple vehicle years
- Fast concurrent page collection
- Uses Dubizzle's Algolia search index directly
- No browser automation required
- No individual listing pages are opened
- Collects up to 150 listings per page
- Automatically discovers the number of available results for each year filter
- Removes duplicate listings
- Streams results directly to the Apify Dataset
- Supports a maximum item limit
- Returns structured data ready for CSV, Excel, JSON, or API usage
- Includes vehicle specifications, seller information, pricing, location, images, and listing links

### How It Works

The Actor uses Dubizzle's search index rather than loading each vehicle page individually.

#### Step 1 — Select the Year Range

Enter the minimum and maximum vehicle years you want to scrape.

For example:

```json
{
    "startYear": 2018,
    "endYear": 2024
}
```

The Actor creates a separate search filter for each requested year.

Older years are grouped into a range where possible to reduce unnecessary requests.

#### Step 2 — Check Available Listings

Before downloading the listings, the Actor sends a lightweight request for each year filter.

This request asks Dubizzle's search index how many matching listings are available.

For example:

```text
2018 → 2,341 listings
2019 → 3,108 listings
2020 → 4,521 listings
```

The Actor then calculates how many pages are required.

#### Step 3 — Download Listing Pages

Each search page contains up to **150 listings**.

The Actor downloads multiple pages concurrently to keep the scraper fast while limiting the number of simultaneous requests.

#### Step 4 — Extract Vehicle Data

Each listing is converted into a structured dataset containing fields such as:

- Year
- Make
- Model
- Kilometers
- Price
- Location
- Seller
- Dealer / End User
- Colour
- Body type
- Fuel type
- Transmission
- Number of cylinders
- Images
- Listing URL
- Regional specifications

#### Step 5 — Remove Duplicates

Duplicate listing URLs are removed before the data is sent to the dataset.

#### Step 6 — Save Results

Results are pushed directly to the **Apify Dataset** while the Actor is running.

You can then export the dataset as:

- JSON
- CSV
- Excel
- XML

or access it through the Apify API.

***

## Input

The Actor accepts the following input parameters.

### Start year

The minimum vehicle year to scrape.

Example:

```json
{
    "startYear": 2020
}
```

Default:

```text
1900
```

### End year

The maximum vehicle year to scrape.

Example:

```json
{
    "endYear": 2025
}
```

The default is the current year.

### Maximum items

Limits the number of listings returned.

Example:

```json
{
    "maxItems": 1000
}
```

Use:

```json
{
    "maxItems": 0
}
```

for unlimited results.

### Maximum concurrency

Controls how many requests can be processed concurrently.

Example:

```json
{
    "maxConcurrency": 5
}
```

The default is **5**.

***

## Example Input

To scrape used cars from 2018 through 2024:

```json
{
    "startYear": 2018,
    "endYear": 2024,
    "maxItems": 0,
    "maxConcurrency": 5
}
```

To collect only 500 listings:

```json
{
    "startYear": 2020,
    "endYear": 2024,
    "maxItems": 500,
    "maxConcurrency": 5
}
```

***

## Output

Each dataset item contains the following fields.

| Field               | Description                        |
| ------------------- | ---------------------------------- |
| `Year`              | Vehicle manufacturing/model year   |
| `Make`              | Vehicle manufacturer               |
| `Model`             | Vehicle model                      |
| `KMs`               | Vehicle mileage                    |
| `Price`             | Listed vehicle price               |
| `Location`          | Listing location                   |
| `Dealer/End User`   | Seller classification              |
| `Date`              | Listing creation date              |
| `Colour`            | Exterior colour                    |
| `Doors/Body`        | Body type or door information      |
| `Title`             | Listing title                      |
| `Link`              | Dubizzle listing URL               |
| `Updated at`        | Reserved field for update tracking |
| `Seller`            | Seller/agent name when available   |
| `Regional Specs`    | Regional specification             |
| `DI`                | Reserved data field                |
| `Image`             | Main listing image                 |
| `Doors`             | Number of doors                    |
| `Body`              | Vehicle body type                  |
| `Fuel Type`         | Fuel type                          |
| `Transmission Type` | Transmission type                  |
| `No. of Cylinders`  | Number of cylinders                |
| `Cruise Control`    | Cruise control information         |

***

## Example Output

```json
{
    "Year": 2022,
    "Make": "Toyota",
    "Model": "Corolla",
    "KMs": "45000",
    "Price": "65000",
    "Location": "Dubai",
    "Dealer/End User": "Dealer",
    "Date": "2026-09-30 14:25:10",
    "Colour": "White",
    "Doors/Body": "Sedan",
    "Title": "Toyota Corolla 2022",
    "Link": "https://uae.dubizzle.com/motors/used-cars/toyota/corolla/...",
    "Updated at": "",
    "Seller": "Example Motors",
    "Regional Specs": "GCC Specs",
    "DI": "",
    "Image": "https://example.com/image.jpg",
    "Doors": "4",
    "Body": "Sedan",
    "Fuel Type": "Petrol",
    "Transmission Type": "Automatic",
    "No. of Cylinders": "4",
    "Cruise Control": "Yes",
}
```

***

## How to Use

#### 1. Open the Actor

Open the **Dubizzle UAE Cars Scraper** Actor.

#### 2. Configure the input

Enter the vehicle year range you need.

For example:

```json
{
    "startYear": 2020,
    "endYear": 2025,
    "maxItems": 1000,
    "maxConcurrency": 5
}
```

#### 3. Start the Actor

Click **Start**.

The Actor will first check how many listings are available for each requested year.

#### 4. Scraping starts automatically

The Actor then downloads the required search pages concurrently.

You can monitor the progress in the Actor log.

Example:

```text
2020 -> nbHits=4521, pages=31
2021 -> nbHits=5182, pages=35
2022 -> nbHits=6210, pages=42
```

#### 5. Open the Dataset

When the run finishes, open the Actor's **Dataset** tab to view the collected cars.

You can export the results or connect the Dataset to another workflow using the Apify API.

***

## Why This Scraper Is Fast

Instead of opening hundreds or thousands of individual vehicle pages, the Actor queries Dubizzle's underlying search index.

A search request can return up to 150 vehicle listings at once.

This significantly reduces the number of HTTP requests required compared with browser-based scraping.

The Actor also processes multiple search pages concurrently.

***

## Data Source

The Actor collects data from the publicly accessible search data used by **Dubizzle UAE**.

It queries the search index directly and does not open individual vehicle detail pages.

The Actor does not require browser automation.

***

## No Browser Required

This Actor does not use:

- Selenium
- Playwright
- Chrome
- Firefox

It uses direct HTTP requests to the search endpoint, making it lightweight and suitable for Apify's Actor environment.

***

## Duplicate Handling

Listings are deduplicated using their listing URL.

If the same listing appears in multiple search results, only one record is added to the Dataset.

***

## Important Notes

Some fields may be empty when Dubizzle does not expose that information in its search index.

In particular, seller names and phone numbers are only available when they are included in the search data.

The Actor does **not** open individual listing pages to retrieve missing fields.

Image URLs are taken from the listing's available image data, with the first available image used as the main image.

***

## Use Cases

This Actor can be used for:

- Used car market research
- Automotive price analysis
- Vehicle inventory monitoring
- Dealer inventory research
- Market intelligence
- Car price comparison
- Automotive lead generation
- Vehicle database creation
- Historical price analysis
- Automotive data pipelines
- Research and analytics

***

## Example Use Case

For example, you can collect all available 2020–2025 used cars:

```json
{
    "startYear": 2020,
    "endYear": 2025,
    "maxItems": 0,
    "maxConcurrency": 5
}
```

The Actor will:

1. Create the required year filters.
2. Check the number of matching listings.
3. Calculate the required number of pages.
4. Download the pages concurrently.
5. Extract the vehicle information.
6. Remove duplicate listings.
7. Push the results directly to the Apify Dataset.

***

## Dataset

The resulting Dataset can be used directly with:

- Apify Dataset export
- CSV
- JSON
- Excel
- Apify API
- Other Apify Actors
- External data pipelines

***

## Support

If you encounter an issue or need additional fields, please check the Actor logs first.

When reporting an issue, include:

- The input used
- The Actor run ID
- The affected year range
- A sample listing URL if applicable

This makes it easier to investigate the issue.

# Actor input Schema

## `startYear` (type: `integer`):

Earliest model year to scrape.

## `endYear` (type: `integer`):

Latest model year to scrape. Leave empty for current year + 1.

## `maxItems` (type: `integer`):

Stop after this many cars. 0 = unlimited.

## `maxConcurrency` (type: `integer`):

Parallel requests to the Algolia API.

## Actor input object example

```json
{
  "startYear": 1900,
  "maxItems": 0,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("husseinelabeedy/dubizzle-uae-vehicles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("husseinelabeedy/dubizzle-uae-vehicles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call husseinelabeedy/dubizzle-uae-vehicles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,husseinelabeedy/dubizzle-uae-vehicles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vZkCEGdEHR7mUJyLp/builds/rTONn70aEn3NXuiiL/openapi.json
