# Booking.com Hotels Scraper (`bgfc97/booking-com-hotels-scraper`) Actor

Scrape Booking.com hotels by destination and dates: name, price per night, review score, stars, address, type and link. Fast, low cost.

- **URL**: https://apify.com/bgfc97/booking-com-hotels-scraper.md
- **Developed by:** [Bruno](https://apify.com/bgfc97) (community)
- **Categories:** Travel
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.70 / 1,000 hotel scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Booking.com Hotels Scraper

Scrape **Booking.com hotel search results** by destination/dates/guests, and/or **individual hotel pages** by URL. Returns name, price, review score & count, address, amenities and more per hotel — using a **real headless browser (Playwright)** behind **Apify residential proxy** to actually pass Booking.com's AWS WAF anti-bot challenge, on both search results and hotel detail pages.

### Why it's useful

Booking.com is one of the most valuable — and most protected — travel data sources. Track hotel prices for a destination and date range, monitor a specific property's rate/review score over time, or build a dataset of hotels for a city.

### Input

```json
{
  "destination": "Lisbon",
  "checkIn": "2026-11-10",
  "checkOut": "2026-11-12",
  "adults": 2,
  "maxItems": 30,
  "currency": "USD"
}
```

or, for specific properties:

```json
{ "hotelUrls": ["https://www.booking.com/hotel/pt/altis-avenida.html"] }
```

- `destination` — place name (e.g. "Lisbon", "Paris, France"). Optional if `hotelUrls` is set.
- `checkIn` / `checkOut` — `YYYY-MM-DD`. Optional; Booking.com uses its own default dates if omitted.
- `adults` — guest count (default 2).
- `hotelUrls` — specific hotel page URLs to scrape directly. Optional if `destination` is set. **At least one of `destination`/`hotelUrls` is required.**
- `maxItems` — cap on hotels returned from the destination search (paginated 25 at a time).
- `currency` — 3-letter code Booking.com should price in.
- `proxyConfiguration` — Apify Proxy settings. **Residential proxies strongly recommended** (see Notes).

### Output

One item per hotel, `source: "search"` or `source: "url"`:

```json
{
  "source": "search",
  "name": "Altis Avenida Hotel",
  "url": "https://www.booking.com/hotel/pt/altis-avenida.html",
  "price": 187,
  "currency": "USD",
  "review_score": 8.9,
  "review_word": "Fabulous",
  "reviews_count": 3421,
  "address": "Lisbon, Portugal",
  "distance": "0.3 km from centre",
  "image": "https://cf.bstatic.com/..."
}
```

Hotel-page items additionally include `description`, `amenities`, `images`, `latitude`/`longitude` and `parsed_with` (`json-ld+meta/dom`, `meta+dom`, or `dom` — which extraction tier actually supplied the fields, so you can judge confidence).

### Honesty on blocking

Every page load runs in a real headless browser behind a residential proxy session. If Booking.com's AWS WAF still serves a challenge/CAPTCHA page, the actor marks that proxy session bad and retries with a **fresh session and a new residential IP** (up to `maxRequestRetries: 4`). Only if every retry is still blocked does it emit `{ ..., error: "BLOCKED_OR_FAILED ..." }` for that item — never fabricated hotel data. If a page loads cleanly but no results can be parsed (no matches, or Booking.com changed its markup), it says so explicitly instead of returning nothing silently.

### Notes / limitations

- Booking.com fronts most pages with an **AWS WAF managed challenge** — a small JS page that must actually run in a browser to pass. This actor uses Playwright (a real headless Chromium) for exactly that reason, unlike a plain-HTTP approach which can spoof headers/TLS but can't execute the challenge JS — that's why a prior HTTP-only version of this actor could load search results (lighter-weight anti-bot) but got hotel *detail* pages blocked outright.
- Search parsing targets the `data-testid` attributes Booking.com's React app renders on result cards (`title`, `price-and-discounted-price`, `review-score`, `address`, `distance`). Hotel-page parsing prefers JSON-LD (`schema.org` Hotel) when present, then Open Graph/meta tags, then the same DOM markers — all reliably-present fields are extracted, nothing is guessed.
- A sign-in popup and/or cookie-consent banner are dismissed automatically if Booking.com shows one, before scraping.

### ⭐ Enjoying this Actor?

A quick **rating/review** helps others find it. Want more fields (amenities, room types, cancellation policy) or multi-page hotel-review scraping added? Open a ticket on the **Issues** tab.

# Actor input Schema

## `destination` (type: `string`):

Place name to search on Booking.com, e.g. "Lisbon" or "Paris, France". Runs a search-results scrape, paginating until maxItems is reached. Optional if hotelUrls is set — at least one of the two is required.

## `checkIn` (type: `string`):

Check-in date for the search, format YYYY-MM-DD (e.g. 2026-11-10). Only used with destination. Leave empty to let Booking.com pick its default dates.

## `checkOut` (type: `string`):

Check-out date for the search, format YYYY-MM-DD (e.g. 2026-11-12). Only used with destination. Leave empty to let Booking.com pick its default dates.

## `adults` (type: `integer`):

Number of adult guests for the search (1-30). Only used with destination.

## `hotelUrls` (type: `array`):

Specific Booking.com hotel page URLs to scrape individually (e.g. https://www.booking.com/hotel/pt/altis-avenida.html). Optional if destination is set — at least one of the two is required.

## `maxItems` (type: `integer`):

Maximum number of hotels to return from the destination search (paginated 25 at a time). Hotel URLs are always scraped individually, counted toward this same total.

## `currency` (type: `string`):

3-letter currency code Booking.com should display prices in (e.g. USD, EUR, GBP).

## `proxyConfiguration` (type: `object`):

Apify Proxy settings used by the headless browser for every page load. RESIDENTIAL is required, not just recommended — Booking.com fronts pages with an AWS WAF challenge that datacenter IPs fail almost every time, even from a real browser.

## `timeoutSecs` (type: `integer`):

How long the headless browser waits for a search/hotel page to render its content before giving up, in seconds (20-180). Also used for browser navigation timeout.

## Actor input object example

```json
{
  "destination": "Lisbon",
  "adults": 2,
  "maxItems": 30,
  "currency": "USD",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "timeoutSecs": 60
}
```

# Actor output Schema

## `dataset` (type: `string`):

Scrape Booking.com hotel search results (by destination, dates, guests) and individual hotel pages (by URL). Returns name, price, review score/count, address, amenities and more. Uses a real headless browser (Playwright) behind Apify residential proxy to pass Booking.com's AWS WAF anti-bot challenge on both search and hotel-detail pages.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "destination": "Lisbon",
    "currency": "USD"
};

// Run the Actor and wait for it to finish
const run = await client.actor("bgfc97/booking-com-hotels-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "destination": "Lisbon",
    "currency": "USD",
}

# Run the Actor and wait for it to finish
run = client.actor("bgfc97/booking-com-hotels-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "destination": "Lisbon",
  "currency": "USD"
}' |
apify call bgfc97/booking-com-hotels-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bgfc97/booking-com-hotels-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xvxCUetSGBzB6VTSt/builds/JKCmyy6M0LN5nRSoJ/openapi.json
