Google Hotels Scraper avatar

Google Hotels Scraper

Pricing

Pay per usage

Go to Apify Store
Google Hotels Scraper

Google Hotels Scraper

Scrapes hotel search results from Google Hotels (google.com/travel/search): name, price per night, rating, review count, star class, coordinates, address and a search link, for given cities/queries or Google Hotels URLs and check-in/check-out dates.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Relay Data Tools

Relay Data Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Categories

Share

Scrapes hotel search results from Google Hotels (google.com/travel/search) for a city/place query or a Google Hotels URL, for given check-in/check-out dates. Returns hotel name, price for the stay, rating, review count, star class, coordinates, a best-effort address, and a search link -- as structured dataset items, no browser required. Pages through further results automatically up to your maxResults.

Verified end to end on the Apify platform: 80 hotels across Paris/Tokyo/New York/Barcelona with the default (datacenter) proxy, 0 errors.

Why this one

The most-used "Google Hotels Scraper" on Apify Store currently has roughly 59% of its runs end in FAILED or TIMED-OUT. This Actor is built around the opposite priority: reliability first. See "How it works" and "Limitations" below for the specific engineering choices that back that up -- plain HTTP instead of a browser, retries with backoff, block/CAPTCHA detection with session rotation, and graceful per-query error items instead of a crashed run.

Who it's for

  • Travel agencies and OTAs doing rate comparison or market scans across many destinations.
  • Revenue managers and hoteliers monitoring their own and competitors' visible pricing on Google.
  • Price-monitoring / rate-parity tools that need a scheduled, structured feed of Google Hotels prices instead of manual lookups.
  • Anyone who wants hotel search results (not just a single property's page) at a given destination and date range, in bulk.

How it works

Google does not offer a public API for google.com/travel/search, but the page embeds its own search results as JSON in the HTML it serves (the same data its JavaScript uses to render the page), inside AF_initDataCallback(...) blocks. This Actor:

  1. Sends a plain HTTP GET to google.com/travel/search per query/URL, with a cookie that skips the EU/EEA cookie-consent interstitial, and checkin=/checkout= query parameters for the dates.
  2. Extracts that embedded JSON and walks it looking for hotel-record-shaped data (structural pattern matching, not fixed array indices -- see src/parser.py), because Google's internal JSON format has no published schema and shifts over time.
  3. If more hotels are needed to reach maxResults, repeats step 1-2 with start=20, start=40, ... (the page size this endpoint appears to use, found empirically -- not documented anywhere), deduplicating hotels across pages by name + coordinates, until maxResults is reached, a page comes back short (Google's signal there's nothing further), two pages in a row add no new hotels, or a safety cap of 8 pages is hit. A later page failing to fetch just stops pagination early and keeps what was already collected -- only a first-page failure fails the whole query/URL.
  4. Applies your filters and pushes one dataset item per hotel.

How currency is applied

Live testing found that Google does not select the displayed price currency from the curr=/gl= query parameters on this endpoint -- repeated requests for the same destination with different curr/gl values, from the same IP, kept returning that destination's local currency regardless (e.g. curr=USD&gl=us against "hotels in Paris" still returned EUR prices). What Google's proxy/geo-targeting documentation describes, and what this Actor relies on instead, is the requester's apparent location: this Actor routes the request through an Apify Proxy exit node in a country associated with your requested currency (e.g. EUR -> Germany, GBP -> United Kingdom -- see CURRENCY_TO_PROXY_COUNTRY in src/main.py), unless you've already set an explicit apifyProxyCountry yourself in proxyConfiguration, in which case your choice is used instead. curr= is still sent on every request (harmless, and may matter for some destinations) but should not be relied on by itself.

This fix has not yet been verified against the live site. It follows directly from the negative result above (varying curr/gl alone did nothing, so geo-routing is what's left) and from Apify Proxy's documented country-targeting behavior, but proxy country selection can't be exercised from a plain local HTTP client (no Apify Proxy credentials in this environment) -- it needs a real run on the Apify platform (e.g. request

currency: "GBP"
with no apifyProxyCountry set and confirm priceCurrency/priceDisplay come back as GBP/£) to confirm it actually changes what Google returns. Regardless of that mechanism, always treat priceCurrency on each item, not the currency you requested, as the source of truth for what a price actually is -- it's read from the same per-hotel structural position Google uses to price that hotel (node[6][1][3] in src/parser.py), not defaulted from your input.

No headless browser is used. This is deliberate: a plain HTTP request is far cheaper (no browser process, no rendering wait, much lower proxy bandwidth) and, in testing, returned the same data a browser would see once the consent cookie is set -- so there was no reliability upside to paying the cost of Playwright/Puppeteer here. If Google changes this page to require JavaScript execution to produce results, this trade-off would need revisiting.

Input

{
"queries": ["hotels in Riga", "hotels in Lisbon"],
"startUrls": [],
"checkIn": "2026-10-19",
"checkOut": "2026-10-21",
"adults": 2,
"currency": "EUR",
"language": "en",
"countryCode": "us",
"maxResults": 30,
"minRating": 4,
"maxPrice": 200,
"hotelClass": [4, 5],
"maxConcurrency": 3,
"proxyConfiguration": { "useApifyProxy": true }
}
FieldTypeDefaultDescription
queriesarray of strings--Free-text searches, e.g. "hotels in Riga". Provide this and/or startUrls.
startUrlsarray of strings--Full Google Hotels URLs, used as-is (their own query string wins).
checkIn / checkOutstring (YYYY-MM-DD)21 / 23 days from todayStay dates, applied to every queries entry (not to startUrls, which carry their own).
adultsinteger2Guests per room.
currencystringUSDRequested currency (ISO 4217). Applied by routing through an Apify Proxy exit node in a matching country (see "How currency is applied"), not by this parameter alone -- the actual currency used for each price is always returned per item in priceCurrency.
languagestringenGoogle UI language (hl).
countryCodestringusGoogle region (gl); influences currency/market.
maxResultsinteger30Cap on hotels returned per query/URL, applied after filters. Values above ~20 trigger automatic pagination (see "How it works") -- up to 8 pages (~160 raw records before dedup/filtering) are fetched per query/URL.
minRating / maxPrice / hotelClassnumber / number / array of integersnoneOptional post-fetch filters. A hotel missing the filtered-on field (e.g. no published star class) is excluded when that filter is set.
maxConcurrencyinteger3How many queries/URLs to fetch in parallel. Kept low on purpose -- see "Rate limiting" below.
proxyConfigurationobjectApify Proxy (default group)Also used to target currency -- see "How currency is applied". Verified working (default datacenter group) on the Apify platform; switch to RESIDENTIAL if you see blocks.

Output (one dataset item per hotel)

{
"hotelName": "Wellton Riga Hotel & SPA",
"starClass": 4,
"rating": 4.3,
"reviewCount": 3706,
"pricePerNight": 60.0,
"priceCurrency": "EUR",
"priceDisplay": "€60",
"latitude": 56.9464117,
"longitude": 24.1139552,
"address": "13.janvāra iela",
"googleHotelsSearchUrl": "https://www.google.com/travel/search?q=Wellton+Riga+Hotel+%26+SPA&checkin=2026-10-19&checkout=2026-10-21",
"searchQuery": "hotels in Riga",
"checkIn": "2026-10-19",
"checkOut": "2026-10-21",
"currency": "EUR",
"scrapedAt": "2026-09-28T18:52:07.109382+00:00",
"error": null
}

If a query/URL cannot be fetched or parsed at all (blocked, timed out, page format unrecognized), the run does not stop -- one item is pushed for it with every hotel field null and error set to a description, instead of the whole Actor crashing.

A ready-to-use Overview table view (hotel, stars, rating, reviews, price, currency, address, query, dates, error) is available in the dataset UI/API.

See samples/sample_output.json for a full real run (39 hotels across "hotels in Riga" and "hotels in Lisbon", 2 nights, ~3 weeks out) and the "Testing" section below for a 60-result paginated run and a multi-page unit test summary.

Rate limiting and blocking

Google will eventually rate-limit or serve a CAPTCHA/consent page to a single IP making many requests. This Actor:

  • Defaults maxConcurrency to 3 and adds a small randomized delay before each request (including between pages of the same paginated query).
  • Retries each request up to 4 times with exponential backoff + jitter.
  • Detects block/CAPTCHA/consent-redirect responses (src/parser.py::detect_block) and treats them as retryable failures, not as "zero hotels found".
  • Rotates the Apify Proxy session on every retry, when a proxyConfiguration is set, so a retry gets a different IP.

For a handful of searches a direct connection (no proxy) is usually fine -- that's how this Actor's parsing/pagination logic was developed and tested. Verified on the Apify platform with the default Apify Proxy group (datacenter): 80 hotels across 4 cities, 0 errors. Set RESIDENTIAL in proxyConfiguration if you see blocks with the default group.

Testing

  • Unit tests: 67 pytest cases (parser, fetcher, pagination, filters, currency-country selection), all offline against real captured/trimmed fixtures or mocked pages -- no live network calls except the one @pytest.mark.network smoke test, which is skipped by default (pytest -m network to run it).
  • Pagination, local, live:
    queries: ["hotels in Paris"], maxResults: 60, currency: "EUR"
    , no proxy (direct connection). Result: 44-46 distinct hotels pushed across multiple pages (varies slightly run to run, as Google's own results do) -- more than a single page's ~20, short of the requested 60 because pagination correctly stopped once Google stopped returning new hotels, not because of a bug. Verified zero duplicate (hotelName, rounded coordinates) keys in the output. This run is also what surfaced and fixed the cross-page coordinate-jitter dedupe bug described in "Limitations".
  • Currency, local, live: confirmed the request always carries curr=<requested> and that priceCurrency is read per-hotel from the page (not defaulted to the input) -- requesting currency: "GBP" from this (EU, non-proxied) environment returned a mix of EUR and GBP in the same run's priceCurrency values (12 EUR / 3 GBP out of 15 hotels for one "hotels in Paris" run), showing currency can genuinely vary per hotel within one response and confirming the per-item extraction is reading real data rather than a page-wide guess. What this local test cannot confirm is whether proxying through a GB-country exit node shifts that mix towards GBP -- see "How currency is applied".
  • Sample output: samples/sample_output.json, regenerated after all fixes in this round -- 39 hotels, "hotels in Riga" + "hotels in Lisbon", 2 nights ~3 weeks out. Field fill rates: hotelName/rating/reviewCount/pricePerNight/priceCurrency/coordinates 100%, starClass 77%, address 3% (see "Limitations" for why address is low and highly destination-dependent).

Limitations

  • This is reverse-engineered from an undocumented internal Google data format, not a published API. It was captured and tested on 2026-09-28 and works as described then, but Google can change this page's structure at any time without notice. The parser is written defensively (structural pattern matching, not fixed indices; every extraction step degrades to null/empty instead of raising) specifically to reduce -- not eliminate -- the blast radius of such a change. A future break would most likely show up as fields going null/an item's error being set, not as a crash.
  • address is best-effort and often null -- confirmed to be a data-availability gap, not an extraction bug. Fill rate varies sharply by destination: for "hotels in Riga" a clean street address (e.g. "Minsterejas iela 8-10") was present for a large share of results across several runs; for "hotels in Paris", a full manual dump of a hotel record that came back null (Hôtel Jardin de Cluny, a real, well-reviewed small hotel -- not an edge case) showed no address text anywhere in it at all -- only coordinates, nearby landmarks/walking times, a rating breakdown, an official website URL, and an unresolvable Google Maps place-ID pair (0x...:0x...) that would need a separate Maps/Places lookup to turn into a real address (out of scope: it would mean one extra HTTP request per hotel, working against this Actor's "cheap single request" design). Across 7 live captures for Paris (140 hotel records total), only 3 had any address text at all (~2%). In short: when Google's own response includes the address, this Actor extracts it; when it doesn't (which for some major cities is most of the time), there's nothing further to extract without adding a second paid request per hotel. Treat null as "not available for this listing," not as an error, and don't expect a high fill rate for large cities.
  • pricePerNight reflects whatever single price Google's page shows for the search you ran (usually the lowest available offer for the stay). It is not guaranteed to always be a strict "per night" figure vs. a stay total for every listing type -- priceDisplay keeps Google's original string alongside it so you can sanity-check.
  • starClass is the official/government star rating where Google publishes one. It is null for hostels, apartments, and other listing types Google mixes into hotel search results, and for some hotels that simply don't have one on file.
  • googleHotelsSearchUrl is a search link, not a guaranteed single-result permalink. Google does not expose a stable per-hotel URL in the data this Actor scrapes; the link is built from the hotel name + your dates and will normally land on/near that hotel, but is not guaranteed to be the only result.
  • Amenity lists and cross-OTA "deal" price comparisons (e.g. Booking.com vs. Expedia vs. the hotel's own site for the same room) are visible in Google's raw data for some listings but were judged too unreliable/inconsistent across record types to expose as a structured field in this version; priceDisplay/pricePerNight reflect one (the lowest found) price only.
  • adults is sent best-effort and was not conclusively verified to change results. checkIn/checkOut were confirmed (live) to change the response; adults was not exhaustively verified the same way. Treat it as "requested" rather than "guaranteed applied".
  • currency/countryCode do not reliably control Google's displayed currency by themselves -- confirmed by live testing, not an assumption. Repeated requests for the same destination with different curr=/gl= values from the same IP kept returning that destination's local currency regardless (e.g. requesting curr=USD&gl=us for "hotels in Paris" still returned EUR). This Actor instead targets currency via Apify Proxy's country selection (see "How currency is applied") -- which has not itself been verified against the live site yet (needs a real platform run with e.g. currency: "GBP" to confirm). Always read the actually-applied currency from each item's priceCurrency, not from the currency you requested.
  • Pagination's page size and stopping behavior are empirical, not documented by Google. start=0/20/40 were observed live to return distinct-but-overlapping hotel sets (union grew 20 -> 30 -> 38 across 3 pages for one query); this may not hold for every destination or change without notice, which is why the loop also stops on a short page or two no-growth pages rather than assuming a fixed total count.
  • Local network testing (parsing, pagination, block-detection) was from a single non-proxied residential IP in the EU. End-to-end platform behavior (proxying, currency targeting, larger-scale concurrency) was separately verified by a real Apify platform run with the default proxy group (80 hotels, 4 cities, 0 errors) -- but that run predates the currency-targeting and pagination changes in this version, so it doesn't cover them.

This Actor only requests the same public search-results page a browser would when you visit Google Hotels and search -- it does not access any account, bypass a paywall, or use credentials. That said, Google's Terms of Service restrict automated access to its properties. You are responsible for using this Actor in line with Google's terms and any law applicable in your jurisdiction; this Actor is provided as a data-extraction tool, not legal advice on whether a particular use is permitted.

FAQ

Why did a query come back with one item and an error instead of hotels? The request was blocked/rate-limited, timed out after retries, or the response didn't match the expected page format. Check the error field's message; consider lowering maxConcurrency or enabling/upgrading proxyConfiguration.

Can I search by a specific date range far in the future? Yes, any checkIn/checkOut Google itself would accept works. Very far-future dates may return fewer or no live prices, the same as on google.com/travel.

Why is priceCurrency on the item different from the currency I requested? Google was found (via live testing) to select the displayed currency from the requester's apparent location, not from the curr=/gl= query parameters -- so this Actor targets currency by routing through an Apify Proxy exit node in a country associated with your requested currency (see "How currency is applied"). If you still see a mismatch: check whether your input already set an explicit apifyProxyCountry (it overrides the currency-based default), and whether your proxy group actually has exit nodes in the target country. priceCurrency on each item always reflects what was actually parsed from the page for that price, which is what you should treat as authoritative regardless.

Why did maxResults: 60 return fewer than 60 hotels for some queries? Pagination stops early if Google stops returning new hotels (a short page, or two pages in a row with no new hotels after dedup) -- some destinations genuinely don't have 60 distinct listings for Google to return. This is expected, not a bug; check the run log for how many pages were fetched.

Does this use a headless browser? No -- see "How it works". This keeps runs fast and cheap, at the cost of depending on Google continuing to embed results as page JSON (see "Limitations").

How is this priced? See PRICING.md for the proposed pay-per-event plan.