# Hostelworld Hostels Scraper (`scrapyx/hostelworld-hostels-scraper`) Actor

Scrapes hostels, guesthouses and budget hotels from Hostelworld by city. Every row carries nightly dorm and private prices, an 8-axis rating breakdown, GPS coordinates, address, district, facilities and images. Optional per-property detail adds policies, house rules and check-in times.

- **URL**: https://apify.com/scrapyx/hostelworld-hostels-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Travel, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hostelworld Hostels Scraper

Scrapes hostels, guesthouses and budget hotels from
**[Hostelworld](https://www.hostelworld.com)** — the world's largest hostel
booking platform — city by city.

Public data only. No login, no cookies, no browser. **No bot challenge of any
kind**: 8 of 8 TLS profiles returned clean 200s on every surface, and the API
answers a request with no headers at all.

This actor reads Hostelworld's own public JSON API (`api.m.hostelworld.com`),
not scraped HTML — so rows arrive already structured, and a site redesign
cannot break them.

### The one thing you need to know before using this

**Hostelworld has no city search endpoint.** Names are resolved through their
own `continents → countries → cities` listing, which means:

| You type | What happens |
| --- | --- |
| `Spain/Barcelona` | **Recommended.** One lookup request, unambiguous. |
| `Barcelona` | Works, but the first bare name builds a global city index (one request per country, run concurrently). |
| `22418` | A raw city id, used as-is with zero lookups. |

That listing is **not exhaustive** — Barcelona is reachable as city `83` too,
and `83` does not appear in it. If a city you know exists is not found, pass
its numeric id instead. A name matching more than one country is refused with
the candidate list rather than guessed at.

Unlike most targets, this API does **not** silently widen a bad input: an
unknown city is a clean `404`, and an unknown property type or sort value is a
clean `400`. What it *does* do is silently ignore parameter names it does not
recognise — so every filter this actor offers was verified to actually move the
result count.

### What you get

Three record types share one dataset, told apart by `recordType`.

#### `PROPERTY` — one row per hostel

Straight from the listing, no detail fetch needed:

- **Prices** — lowest nightly rate overall, and split by dorm vs private room,
  with currency
- **Ratings** — overall score and review count, plus an 8-axis breakdown
  (security, location, staff, atmosphere, cleanliness, facilities, value, fun)
- **Location** — street address, district, GPS coordinates, distance
- **Everything else** — property type, star rating, images, gallery, overview
  text, free-cancellation status, promoted/featured/new flags

With **Fetch full property details** on, each row also gets a `detail` object:
full description, house rules and things to note, cancellation policy, payment
methods, check-in/check-out window, the categorised facility list and a sample
of recent reviews.

> `detail` is kept as a nested object rather than merged into the row on
> purpose — the detail document reuses several key names with a *different*
> type (`id` is an int in the listing and a string in the detail), so a flat
> merge would make column types unstable mid-dataset.

#### `SEARCH_SUMMARY` — one row per city

Hostelworld's own match total, which city id the name resolved to, the city and
country it actually served, how many requests were spent, and `locationApplied`.

#### `ERROR` — one row per input that could not be processed

Every input maps to at least one row, so nothing disappears silently.

### Input

| Field | What it does |
| --- | --- |
| **Cities** | One city per entry — `Spain/Barcelona`, `Barcelona`, or `22418`. |
| **Property type** | `HOSTEL`, `HOTEL`, `GUESTHOUSE`, `APARTMENT`, `CAMPSITE`. Verified to partition a city exactly (Barcelona: 93+5+15+2 = 115). |
| **Sort order** | Rating, price, distance or name. Values Hostelworld refuses (`popularity`, `recommended`) are not offered. |
| **Check-in date + Nights** | Set together to keep only properties with real availability — on London this narrowed 94 to 80. One without the other is an upstream `400`, so they are gated as a pair. |
| **Fetch full property details** | Adds one request per property. Off by default. |
| **Max properties per city** | `0` for unlimited. There is no result-window ceiling. |
| **Rows per request** | Up to 500. No server-side cap was found — a whole city fits in one request. |
| **Max concurrent requests / Minimum seconds between requests** | Throughput controls. No throttling was observed, so the interval defaults to `0`. |

### Known limits

- **Cities, not regions or countries.** The API is organised per city; there is
  no "all of Spain" query. Pass several cities instead.
- **The country city listing is incomplete.** See above — numeric ids are the
  escape hatch.
- **Prices are indicative.** Listing prices are Hostelworld's "from" rates. For
  rates tied to specific dates, set check-in and nights.
- **Reviews are not included.** `/properties/{id}/reviews` is a separate
  paginated endpoint — one Barcelona hostel alone has 698 reviews at 50 per
  page. Folding that in would turn a 100-property crawl into thousands of
  requests, so it belongs in a sibling actor.
- **Availability is a filter, not a room inventory.** Check-in + nights narrows
  *which properties* are returned; it does not return per-room availability.

### Notes on politeness

The API showed no rate limiting (12 back-to-back detail calls all returned
200\), so defaults are modest rather than aggressive: concurrency 4, no forced
delay. `Minimum seconds between requests` paces request *starts* without
occupying a concurrency slot, so raising it slows the crawl without wasting
workers.

# Actor input Schema

## `locations` (type: `array`):

One city per entry, each with its own SEARCH\_SUMMARY row.

Three accepted forms:

- `Spain/Barcelona` — **recommended**. Country-scoped, unambiguous, and needs only one lookup request.
- `Barcelona` — bare name. Works, but the first bare name in a run builds a global city index (one request per country, issued concurrently). Refused with a list of candidates if the name exists in more than one country.
- `22418` — a raw Hostelworld city id, used as-is with no lookup.

Hostelworld has no city-search endpoint, so names are resolved through their own continents → countries → cities listing. That listing is not exhaustive: if a city you know exists is not found, pass its numeric id instead.

## `propertyType` (type: `string`):

Restrict to one kind of property. Verified on Barcelona, where the types partition the city exactly: 93 hostels + 5 hotels + 15 guesthouses + 2 apartments = the full 115. An unrecognised type is refused by Hostelworld with HTTP 400 rather than being silently ignored, so anything outside this list is rejected before the run starts.

## `sort` (type: `string`):

Each value here was verified to genuinely reorder results. `popularity` and `recommended` look plausible but are refused with HTTP 400, so they are not offered.

## `checkInDate` (type: `string`):

Optional `YYYY-MM-DD`. Set this together with **Nights** to return only properties with real availability for those dates — on London this narrowed 94 properties to 80. Must be set as a pair: one without the other is answered with HTTP 400.

## `nights` (type: `integer`):

Length of stay in nights, used with **Check-in date**. Leave at 0 to list every property regardless of availability.

## `includePropertyDetails` (type: `boolean`):

Adds one request per property and attaches a `detail` object carrying what the listing does not: full description, house rules and things to note, cancellation policy, payment methods, check-in/check-out window, the categorised facility list and a sample of recent reviews.

Off by default — search rows already include prices, the 8-axis rating breakdown, coordinates, address, district and images, which is enough for most uses.

## `maxItems` (type: `integer`):

Stop after this many properties per city. Set to 0 for unlimited — Hostelworld imposes no result-window ceiling, and the largest city measured (Bangkok) holds 326 properties.

## `pageSize` (type: `integer`):

How many properties to pull per API call. No server-side cap was found — 500 returns a whole city in a single request. 200 keeps responses around 1 MB while still covering most cities in one call.

## `maxConcurrency` (type: `integer`):

Upper bound on requests in flight at once, across all cities and any detail fetches.

## `minRequestInterval` (type: `integer`):

Paces how often requests START, without tying up a concurrency slot. No throttling was observed (12 back-to-back detail calls all returned 200), so this defaults to 0 — raise it to be gentler on large crawls.

## `proxyConfiguration` (type: `object`):

Hostelworld's API is fully open — no auth, no token, and no bot challenge on any of the 8 TLS profiles tested across every surface. Residential is still the cloud default, because container egress is a different posture than a home connection. No country pin: this is a global travel API with no geo gate.

## Actor input object example

```json
{
  "locations": [
    "Spain/Barcelona"
  ],
  "propertyType": "",
  "sort": "",
  "checkInDate": "",
  "nights": 0,
  "includePropertyDetails": false,
  "maxItems": 100,
  "pageSize": 200,
  "maxConcurrency": 4,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "locations": [
        "Spain/Barcelona"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/hostelworld-hostels-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "locations": ["Spain/Barcelona"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/hostelworld-hostels-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "locations": [
    "Spain/Barcelona"
  ]
}' |
apify call scrapyx/hostelworld-hostels-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/hostelworld-hostels-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vpjiULo2scZeM3aIW/builds/mFhEf5pTuc0jIw7Kj/openapi.json
