Google Hotels Scraper
Pricing
Pay per usage
Google Hotels Scraper
Scrapes hotel search results from Google Hotels (google.com/travel/search): name, price per night, rating, review count, star class, coordinates, address and a search link, for given cities/queries or Google Hotels URLs and check-in/check-out dates.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Relay Data Tools
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
Scrapes hotel search results from Google Hotels (google.com/travel/search) for a
city/place query or a Google Hotels URL, for given check-in/check-out dates. Returns hotel
name, price for the stay, rating, review count, star class, coordinates, a best-effort
address, and a search link -- as structured dataset items, no browser required. Pages
through further results automatically up to your maxResults.
Verified end to end on the Apify platform: 80 hotels across Paris/Tokyo/New York/Barcelona with the default (datacenter) proxy, 0 errors.
Why this one
The most-used "Google Hotels Scraper" on Apify Store currently has roughly 59% of its runs end in FAILED or TIMED-OUT. This Actor is built around the opposite priority: reliability first. See "How it works" and "Limitations" below for the specific engineering choices that back that up -- plain HTTP instead of a browser, retries with backoff, block/CAPTCHA detection with session rotation, and graceful per-query error items instead of a crashed run.
Who it's for
- Travel agencies and OTAs doing rate comparison or market scans across many destinations.
- Revenue managers and hoteliers monitoring their own and competitors' visible pricing on Google.
- Price-monitoring / rate-parity tools that need a scheduled, structured feed of Google Hotels prices instead of manual lookups.
- Anyone who wants hotel search results (not just a single property's page) at a given destination and date range, in bulk.
How it works
Google does not offer a public API for google.com/travel/search, but the page embeds its
own search results as JSON in the HTML it serves (the same data its JavaScript uses to
render the page), inside AF_initDataCallback(...) blocks. This Actor:
- Sends a plain HTTP GET to
google.com/travel/searchper query/URL, with a cookie that skips the EU/EEA cookie-consent interstitial, andcheckin=/checkout=query parameters for the dates. - Extracts that embedded JSON and walks it looking for hotel-record-shaped data
(structural pattern matching, not fixed array indices -- see
src/parser.py), because Google's internal JSON format has no published schema and shifts over time. - If more hotels are needed to reach
maxResults, repeats step 1-2 withstart=20,start=40, ... (the page size this endpoint appears to use, found empirically -- not documented anywhere), deduplicating hotels across pages by name + coordinates, untilmaxResultsis reached, a page comes back short (Google's signal there's nothing further), two pages in a row add no new hotels, or a safety cap of 8 pages is hit. A later page failing to fetch just stops pagination early and keeps what was already collected -- only a first-page failure fails the whole query/URL. - Applies your filters and pushes one dataset item per hotel.
How currency is applied
Live testing found that Google does not select the displayed price currency from the
curr=/gl= query parameters on this endpoint -- repeated requests for the same
destination with different curr/gl values, from the same IP, kept returning that
destination's local currency regardless (e.g. curr=USD&gl=us against "hotels in Paris"
still returned EUR prices). What Google's proxy/geo-targeting documentation describes, and
what this Actor relies on instead, is the requester's apparent location: this Actor routes
the request through an Apify Proxy exit node in a country associated with your requested
currency (e.g. EUR -> Germany, GBP -> United Kingdom -- see
CURRENCY_TO_PROXY_COUNTRY in src/main.py), unless you've already set an explicit
apifyProxyCountry yourself in proxyConfiguration, in which case your choice is used
instead. curr= is still sent on every request (harmless, and may matter for some
destinations) but should not be relied on by itself.
This fix has not yet been verified against the live site. It follows directly from the
negative result above (varying curr/gl alone did nothing, so geo-routing is what's
left) and from Apify Proxy's documented country-targeting behavior, but proxy country
selection can't be exercised from a plain local HTTP client (no Apify Proxy credentials in
this environment) -- it needs a real run on the Apify platform (e.g. request
currency: "GBP"apifyProxyCountry set and confirm priceCurrency/priceDisplay come back
as GBP/£) to confirm it actually changes what Google returns. Regardless of that mechanism,
always treat priceCurrency on each item, not the currency you requested, as the source
of truth for what a price actually is -- it's read from the same per-hotel structural
position Google uses to price that hotel (node[6][1][3] in src/parser.py), not
defaulted from your input.
No headless browser is used. This is deliberate: a plain HTTP request is far cheaper (no browser process, no rendering wait, much lower proxy bandwidth) and, in testing, returned the same data a browser would see once the consent cookie is set -- so there was no reliability upside to paying the cost of Playwright/Puppeteer here. If Google changes this page to require JavaScript execution to produce results, this trade-off would need revisiting.
Input
{"queries": ["hotels in Riga", "hotels in Lisbon"],"startUrls": [],"checkIn": "2026-10-19","checkOut": "2026-10-21","adults": 2,"currency": "EUR","language": "en","countryCode": "us","maxResults": 30,"minRating": 4,"maxPrice": 200,"hotelClass": [4, 5],"maxConcurrency": 3,"proxyConfiguration": { "useApifyProxy": true }}
| Field | Type | Default | Description |
|---|---|---|---|
queries | array of strings | -- | Free-text searches, e.g. "hotels in Riga". Provide this and/or startUrls. |
startUrls | array of strings | -- | Full Google Hotels URLs, used as-is (their own query string wins). |
checkIn / checkOut | string (YYYY-MM-DD) | 21 / 23 days from today | Stay dates, applied to every queries entry (not to startUrls, which carry their own). |
adults | integer | 2 | Guests per room. |
currency | string | USD | Requested currency (ISO 4217). Applied by routing through an Apify Proxy exit node in a matching country (see "How currency is applied"), not by this parameter alone -- the actual currency used for each price is always returned per item in priceCurrency. |
language | string | en | Google UI language (hl). |
countryCode | string | us | Google region (gl); influences currency/market. |
maxResults | integer | 30 | Cap on hotels returned per query/URL, applied after filters. Values above ~20 trigger automatic pagination (see "How it works") -- up to 8 pages (~160 raw records before dedup/filtering) are fetched per query/URL. |
minRating / maxPrice / hotelClass | number / number / array of integers | none | Optional post-fetch filters. A hotel missing the filtered-on field (e.g. no published star class) is excluded when that filter is set. |
maxConcurrency | integer | 3 | How many queries/URLs to fetch in parallel. Kept low on purpose -- see "Rate limiting" below. |
proxyConfiguration | object | Apify Proxy (default group) | Also used to target currency -- see "How currency is applied". Verified working (default datacenter group) on the Apify platform; switch to RESIDENTIAL if you see blocks. |
Output (one dataset item per hotel)
{"hotelName": "Wellton Riga Hotel & SPA","starClass": 4,"rating": 4.3,"reviewCount": 3706,"pricePerNight": 60.0,"priceCurrency": "EUR","priceDisplay": "€60","latitude": 56.9464117,"longitude": 24.1139552,"address": "13.janvāra iela","googleHotelsSearchUrl": "https://www.google.com/travel/search?q=Wellton+Riga+Hotel+%26+SPA&checkin=2026-10-19&checkout=2026-10-21","searchQuery": "hotels in Riga","checkIn": "2026-10-19","checkOut": "2026-10-21","currency": "EUR","scrapedAt": "2026-09-28T18:52:07.109382+00:00","error": null}
If a query/URL cannot be fetched or parsed at all (blocked, timed out, page format
unrecognized), the run does not stop -- one item is pushed for it with every hotel field
null and error set to a description, instead of the whole Actor crashing.
A ready-to-use Overview table view (hotel, stars, rating, reviews, price, currency, address, query, dates, error) is available in the dataset UI/API.
See samples/sample_output.json for a full real run (39 hotels across "hotels in Riga" and
"hotels in Lisbon", 2 nights, ~3 weeks out) and the "Testing" section below for a 60-result
paginated run and a multi-page unit test summary.
Rate limiting and blocking
Google will eventually rate-limit or serve a CAPTCHA/consent page to a single IP making many requests. This Actor:
- Defaults
maxConcurrencyto 3 and adds a small randomized delay before each request (including between pages of the same paginated query). - Retries each request up to 4 times with exponential backoff + jitter.
- Detects block/CAPTCHA/consent-redirect responses (
src/parser.py::detect_block) and treats them as retryable failures, not as "zero hotels found". - Rotates the Apify Proxy session on every retry, when a
proxyConfigurationis set, so a retry gets a different IP.
For a handful of searches a direct connection (no proxy) is usually fine -- that's how this
Actor's parsing/pagination logic was developed and tested. Verified on the Apify platform
with the default Apify Proxy group (datacenter): 80 hotels across 4 cities, 0 errors. Set
RESIDENTIAL in proxyConfiguration if you see blocks with the default group.
Testing
- Unit tests: 67 pytest cases (parser, fetcher, pagination, filters, currency-country
selection), all offline against real captured/trimmed fixtures or mocked pages -- no live
network calls except the one
@pytest.mark.networksmoke test, which is skipped by default (pytest -m networkto run it). - Pagination, local, live: , no proxy (direct connection). Result: 44-46 distinct hotels pushed across multiple pages (varies slightly run to run, as Google's own results do) -- more than a single page's ~20, short of the requested 60 because pagination correctly stopped once Google stopped returning new hotels, not because of a bug. Verified zero duplicatequeries: ["hotels in Paris"], maxResults: 60, currency: "EUR"
(hotelName, rounded coordinates)keys in the output. This run is also what surfaced and fixed the cross-page coordinate-jitter dedupe bug described in "Limitations". - Currency, local, live: confirmed the request always carries
curr=<requested>and thatpriceCurrencyis read per-hotel from the page (not defaulted to the input) -- requestingcurrency: "GBP"from this (EU, non-proxied) environment returned a mix ofEURandGBPin the same run'spriceCurrencyvalues (12 EUR / 3 GBP out of 15 hotels for one "hotels in Paris" run), showing currency can genuinely vary per hotel within one response and confirming the per-item extraction is reading real data rather than a page-wide guess. What this local test cannot confirm is whether proxying through a GB-country exit node shifts that mix towards GBP -- see "How currency is applied". - Sample output:
samples/sample_output.json, regenerated after all fixes in this round -- 39 hotels, "hotels in Riga" + "hotels in Lisbon", 2 nights ~3 weeks out. Field fill rates: hotelName/rating/reviewCount/pricePerNight/priceCurrency/coordinates 100%, starClass 77%, address 3% (see "Limitations" for why address is low and highly destination-dependent).
Limitations
- This is reverse-engineered from an undocumented internal Google data format, not a
published API. It was captured and tested on 2026-09-28 and works as described then, but
Google can change this page's structure at any time without notice. The parser is written
defensively (structural pattern matching, not fixed indices; every extraction step
degrades to
null/empty instead of raising) specifically to reduce -- not eliminate -- the blast radius of such a change. A future break would most likely show up as fields goingnull/an item'serrorbeing set, not as a crash. addressis best-effort and oftennull-- confirmed to be a data-availability gap, not an extraction bug. Fill rate varies sharply by destination: for "hotels in Riga" a clean street address (e.g."Minsterejas iela 8-10") was present for a large share of results across several runs; for "hotels in Paris", a full manual dump of a hotel record that came backnull(Hôtel Jardin de Cluny, a real, well-reviewed small hotel -- not an edge case) showed no address text anywhere in it at all -- only coordinates, nearby landmarks/walking times, a rating breakdown, an official website URL, and an unresolvable Google Maps place-ID pair (0x...:0x...) that would need a separate Maps/Places lookup to turn into a real address (out of scope: it would mean one extra HTTP request per hotel, working against this Actor's "cheap single request" design). Across 7 live captures for Paris (140 hotel records total), only 3 had any address text at all (~2%). In short: when Google's own response includes the address, this Actor extracts it; when it doesn't (which for some major cities is most of the time), there's nothing further to extract without adding a second paid request per hotel. Treatnullas "not available for this listing," not as an error, and don't expect a high fill rate for large cities.pricePerNightreflects whatever single price Google's page shows for the search you ran (usually the lowest available offer for the stay). It is not guaranteed to always be a strict "per night" figure vs. a stay total for every listing type --priceDisplaykeeps Google's original string alongside it so you can sanity-check.starClassis the official/government star rating where Google publishes one. It isnullfor hostels, apartments, and other listing types Google mixes into hotel search results, and for some hotels that simply don't have one on file.googleHotelsSearchUrlis a search link, not a guaranteed single-result permalink. Google does not expose a stable per-hotel URL in the data this Actor scrapes; the link is built from the hotel name + your dates and will normally land on/near that hotel, but is not guaranteed to be the only result.- Amenity lists and cross-OTA "deal" price comparisons (e.g. Booking.com vs. Expedia
vs. the hotel's own site for the same room) are visible in Google's raw data for some
listings but were judged too unreliable/inconsistent across record types to expose as a
structured field in this version;
priceDisplay/pricePerNightreflect one (the lowest found) price only. adultsis sent best-effort and was not conclusively verified to change results.checkIn/checkOutwere confirmed (live) to change the response;adultswas not exhaustively verified the same way. Treat it as "requested" rather than "guaranteed applied".currency/countryCodedo not reliably control Google's displayed currency by themselves -- confirmed by live testing, not an assumption. Repeated requests for the same destination with differentcurr=/gl=values from the same IP kept returning that destination's local currency regardless (e.g. requestingcurr=USD&gl=usfor "hotels in Paris" still returned EUR). This Actor instead targets currency via Apify Proxy's country selection (see "How currency is applied") -- which has not itself been verified against the live site yet (needs a real platform run with e.g.currency: "GBP"to confirm). Always read the actually-applied currency from each item'spriceCurrency, not from thecurrencyyou requested.- Pagination's page size and stopping behavior are empirical, not documented by
Google.
start=0/20/40were observed live to return distinct-but-overlapping hotel sets (union grew 20 -> 30 -> 38 across 3 pages for one query); this may not hold for every destination or change without notice, which is why the loop also stops on a short page or two no-growth pages rather than assuming a fixed total count. - Local network testing (parsing, pagination, block-detection) was from a single non-proxied residential IP in the EU. End-to-end platform behavior (proxying, currency targeting, larger-scale concurrency) was separately verified by a real Apify platform run with the default proxy group (80 hotels, 4 cities, 0 errors) -- but that run predates the currency-targeting and pagination changes in this version, so it doesn't cover them.
Legal
This Actor only requests the same public search-results page a browser would when you visit Google Hotels and search -- it does not access any account, bypass a paywall, or use credentials. That said, Google's Terms of Service restrict automated access to its properties. You are responsible for using this Actor in line with Google's terms and any law applicable in your jurisdiction; this Actor is provided as a data-extraction tool, not legal advice on whether a particular use is permitted.
FAQ
Why did a query come back with one item and an error instead of hotels?
The request was blocked/rate-limited, timed out after retries, or the response didn't match
the expected page format. Check the error field's message; consider lowering
maxConcurrency or enabling/upgrading proxyConfiguration.
Can I search by a specific date range far in the future?
Yes, any checkIn/checkOut Google itself would accept works. Very far-future dates may
return fewer or no live prices, the same as on google.com/travel.
Why is priceCurrency on the item different from the currency I requested?
Google was found (via live testing) to select the displayed currency from the requester's
apparent location, not from the curr=/gl= query parameters -- so this Actor targets
currency by routing through an Apify Proxy exit node in a country associated with your
requested currency (see "How currency is applied"). If you still see a mismatch: check
whether your input already set an explicit apifyProxyCountry (it overrides the
currency-based default), and whether your proxy group actually has exit nodes in the target
country. priceCurrency on each item always reflects what was actually parsed from the
page for that price, which is what you should treat as authoritative regardless.
Why did maxResults: 60 return fewer than 60 hotels for some queries?
Pagination stops early if Google stops returning new hotels (a short page, or two pages in
a row with no new hotels after dedup) -- some destinations genuinely don't have 60 distinct
listings for Google to return. This is expected, not a bug; check the run log for how many
pages were fetched.
Does this use a headless browser? No -- see "How it works". This keeps runs fast and cheap, at the cost of depending on Google continuing to embed results as page JSON (see "Limitations").
How is this priced? See PRICING.md for the proposed pay-per-event plan.