Google Search Results Scraper avatar

Google Search Results Scraper

Pricing

from $1.00 / 1,000 results

Go to Apify Store
Google Search Results Scraper

Google Search Results Scraper

Scrape Google Search result pages (SERPs) and extract structured data: organic results, paid ads, related queries, and People Also Ask. Supports country/language targeting, time filters, pagination, and CSV-friendly output.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Crawler Bros

Crawler Bros

Maintained by Community

Actor stats

0

Bookmarked

95

Total users

11

Monthly active users

a day ago

Last modified

Share

Google Search Results (SERP) Scraper

Scrape Google Search result pages and get structured data from any query. Extract organic results, AI Overviews, People Also Ask, local packs, top stories, and more — all in clean JSON or CSV format.

What does Google Search Results Scraper do?

This scraper lets you search Google at scale and collect structured data from the results. Just enter your search queries and get back organized results including:

  • Organic search results (title, URL, description, position, sitelinks)
  • Google AI Overview text and cited sources (when Google shows one)
  • Knowledge panel, People Also Ask, related searches, local pack, and top stories
  • Sports standings tables and result counts / pagination info
  • Optional weather, time-zone, jobs, video, and shopping pack info boxes

It works across 48 countries and supports language targeting, time filters, custom date ranges, precise city-level geo-targeting, exact-phrase matching, search operators (site/intitle/intext/inurl/filetype/exclude terms), mobile results, and multi-page pagination.

How to use

  1. Enter your search queries — Add one or more search terms
  2. Choose a country — Select from 48 countries (US by default)
  3. Set filters — Optionally filter by time period, language, date range, or search operators
  4. Run the scraper — Results appear in the dataset within seconds

No login or API key needed. The scraper handles proxy rotation and rate limiting automatically, and automatically adapts if it detects it's being blocked, so you get reliable results without any manual tuning.

Input options

OptionDescriptionDefault
Search QueriesList of queries to searchRequired
CountryTarget country (determines Google domain and the gl searcher-location signal, and steers the automatic proxy chain toward an exit IP in that country). If left empty and Precise Location (uule) is set, the country is instead derived from the uule valueUnited States
Restrict Results Country (cr)Only return pages originating from/targeting this country - independent of Country above, which only sets the searcher's locationAny country
LanguageLanguage of the Google interface/UI itself (e.g., en, de, fr) — does not restrict which language the matched pages are inDefault
Restrict Results Language (lr)Only return results for pages written in this language, independent of Language above. Comma-separate multiple codes to match any of them, e.g. en,frAny language
Max ResultsMaximum organic results per query (1-200)100
Max PagesPages to crawl per query (1-10)1
Results Per PageResults per page — snapped to Google's supported values (10, 20, 30, 40, 50, 100). Google has largely stopped honoring this for organic results since late 2023 and typically still returns roughly 7-10 organic results per page fetch no matter how high this is set — use Max Pages to fetch more results instead10
Time PeriodFilter: any time, past hour/day/week/month/yearAny time
SafeSearchFilter explicit content: default, active (filter on), or offDefault
Date From / Date ToCustom date range (YYYY-MM-DD). Must be set together — a lone value is ignoredOff
Precise Location (uule)Google-encoded location string for city-level geo-targeting. Its country is automatically used to steer the default proxy chain and the gl signal (no need to also set Country separately for the same place) — but city-level precision beyond the country match is still best-effort: Google can fall back to the proxy IP's real location for the specific city, so always spot-check resultsOff
Exact Phrase MatchWrap each query in quotes so Google matches it as an exact phrase (e.g. "running shoes") instead of matching the words independently and in any orderOff
Site (site:)Restrict results to a domain, e.g. nike.com. Verified — see note belowOff
Word in Title (intitle:)Search operator. Sent to Google as-is, not verified against returned results (see note below)Off
Word in Text (intext:)Search operator. Sent to Google as-is, not verified against returned results (see note below)Off
Word in URL (inurl:)Search operator. Sent to Google as-is, not verified against returned results (see note below)Off
File Type (filetype:)Search operator, e.g. pdf, doc, xls. Sent to Google as-is, not verified against returned results (see note below)Off
Exclude Terms (-word)Words to exclude from results, e.g. jaguar to drop pages matching that word. Separate multiple words with a comma or spaceOff
Include Extra Info BoxesAlso extract weather, time-zone, currency-converter, unit-converter, stock, jobs-pack, video-pack, shopping-pack, and dictionary boxes when present (see Tier 2 fields below)Off
Force AI Overview CaptureForces a slower, JavaScript-rendering mode and waits for Google's AI Overview to load, so aiOverview can populate (slower, best-effort)Off
Mobile ResultsGet mobile version of search results (experimental for organic results — see Mobile support)Off
Proxy ConfigurationStandard Apify Proxy configuration. Leave it unset and the scraper works out of the box: it tries the Google SERP proxy group first, then RESIDENTIAL, then Apify's Web Unblocker, then your account default — the Google SERP/RESIDENTIAL attempts are automatically steered to an exit IP in your Country (or the country derived from uule) rather than left unsteered. Set this explicitly yourself and that automatic steering is skipped, so add your own apifyProxyCountry if geo-targeting mattersGoogle SERP proxy
CSV Friendly OutputOne result per row (for CSV/Excel export)Off

Multi-word values for the site/intitle/intext/inurl/filetype operators are automatically wrapped in quotes so Google scopes the whole phrase to the operator instead of treating the extra words as separate search terms.

For Exclude Terms, each word is applied as its own -word exclusion, e.g. entering jaguar, cats excludes results containing "jaguar" and results containing "cats".

Output example

Each run produces a dataset with structured results. Here's what a typical result looks like:

{
"searchQuery": {
"term": "Hotels in NYC",
"page": 1,
"countryCode": "us",
"languageCode": "en",
"resultsPerPage": 10
},
"hasNextPage": true,
"resultsTotal": 136000000,
"organicResults": [
{
"position": 1,
"title": "THE 10 BEST Hotels in New York City",
"url": "https://www.google.com/goto?url=<opaque-token>",
"displayedUrl": "https://www.tripadvisor.com › Hotels",
"snippet": "Some of the best hotels in New York City are...",
"siteLinks": []
}
],
"aiOverview": null,
"knowledgePanel": null,
"peopleAlsoAsk": [
{ "question": "What is the best area to stay in NYC?" }
],
"relatedQueries": [
{ "query": "Cheap hotels in NYC", "url": "https://www.google.com/search?q=Cheap+hotels+in+NYC" }
],
"localPack": null,
"topStories": null,
"sportsTable": null,
"paidResults": null
}

Output fields

Tier 1 fields (always attempted)

These fields are extracted on every request. Most are null/absent when Google simply doesn't render that feature for a given query — that's expected behavior, not a bug.

FieldDescription
organicResultsTitle, URL, displayed URL, snippet, position, and sitelinks for each organic result. Important: url is Google's own redirect link in the form https://www.google.com/goto?url=<opaque token> — the token is an opaque encoded blob, not the destination URL in any decodable form, so it cannot be parsed or regexed out of the field. Clicking the link does take you to the real result, but recovering the destination programmatically would require following the redirect over the network (an extra round-trip per result), which this actor doesn't currently do. Use displayedUrl instead — it's recovered separately from the visible citation text and reflects the real result domain for the vast majority of results. Exception: for social-media citation cards (e.g. an Instagram-style organic result), Google renders the citation area with follower/like-count text instead of the domain (e.g. Arabic "أكثر من ١٣٦ ألف متابع", Italian "Oltre 10 Mi piace · 2 anni fa"); the actor detects this and falls back to the handle shown elsewhere on the same card (e.g. khobar.rest) when available, otherwise displayedUrl surfaces that engagement text as a best-effort value rather than the domain.
aiOverviewText and cited sources for Google's AI-generated summary. Only extractable when the scraper is running in its JavaScript-rendering mode for that request. Google injects the container this field is read from via client-side JavaScript after the page loads, so it is never present in the actor's default, faster fetch mode — even for queries where Google genuinely rendered an AI Overview. The actor only switches to the slower mode after repeated blocks, so on a healthy default run aiOverview will typically be null regardless of whether Google showed one for the query. Enable "Force AI Overview capture" (forceAiOverviewCapture) to request it (best-effort): this forces every fetch into JavaScript-rendering mode and waits specifically for the AI Overview to finish loading before reading the page. It's slower and more resource-intensive per request, and it's still not a 100% guarantee — Google may genuinely not render an AI Overview for a given query, in which case aiOverview stays null even with the flag on. See the FAQ below for an important tip on proxy configuration with this flag.
knowledgePanelEntity info panel (people, places, organizations). Present only when Google shows one for the query.
peopleAlsoAsk"People also ask" questions. Answers are not extracted — Google loads them on click rather than including them in the static HTML, so only the question text is available. Present only when Google shows this box.
relatedQueriesRelated search suggestions shown at the bottom of the results page.
localPackLocal business/map results block. Present only for local-intent queries. rating, reviewCount, and address are parsed structurally from Google's markup rather than by matching English-phrased text, so parsing itself doesn't depend on the UI language — this works correctly across comma-decimal ratings (e.g. German "4,5"), Eastern/Extended Arabic-Indic digit review counts (e.g. Arabic "٣٬٦٧٩"), and word-based review-count magnitude suffixes in Arabic, Turkish, Korean, Russian, Portuguese (Brazil), Vietnamese, Greek, and Indonesian, plus Hebrew's RTL-marked review-count/address rendering. Any of the three fields can still be independently null if a given business simply doesn't have that data on the page (e.g. no rating yet).
topStoriesNews carousel results. Present only when Google shows a Top Stories box.
sportsTableTitle, headers, and rows for sports-standings-style queries.
paidResultsAd results (#tads/#tadsb slots). Best-effort and often empty: Google frequently renders zero ads for a given query — even commercial/ad-heavy search terms — so null here is common and does not indicate a scraping failure. Both absolute ad links and Google's relative /goto?url=<token> redirect links are recognized; as with organicResults, a redirect-style url points at Google's mediated link rather than the advertiser's domain.
resultsTotalTotal result count, parsed structurally from the digits in Google's result-count text (#result-stats), populating the same way across UI languages (e.g. German "Ungefähr 104 Ergebnisse", Japanese "104 件の結果", Arabic "حوالي 15,900,000 نتيجة"). The extractor reads both a real <div id="result-stats">...</div> DOM node and, when Google instead delivers that markup only as an escaped string inside an inline <script> hydration payload, the script-embedded form — so this field is reliably populated on the vast majority of requests. It can still be null on rare pages that omit the result-count text entirely (e.g. some zero-result or heavily-filtered queries). This number is Google's own rough estimate, not a stable per-query total — it can swing dramatically between consecutive pages of the identical query within a single run. Don't treat it as a fixed count when using maxPagesPerQuery > 1 — expect it to fluctuate page-to-page for the same query, since Google recomputes the estimate independently per page.
hasNextPageReliable boolean signal for whether another results page is available.

Tier 2 fields (opt-in, via includeExtraBoxes)

Off by default to keep typical runs lean. Enable Include Extra Info Boxes to extract these:

FieldDescription
weatherBoxWeather info box, populated when Google renders one for the query. Always includes rawText (the box's full text) plus structured sub-fields: temperature ({value, unit}), condition (e.g. "Sunny", "Cloudy"), precipitationPercent, humidityPercent, and wind ({speed, unit}). location is always null: Google's weather box never includes the resolved location in its own markup (it's inferred purely from the query), so this actor doesn't guess it from the query string. Any structured sub-field can independently be null if Google's markup for that particular result doesn't match the expected shape — rawText is always the reliable fallback.
timeZoneBoxTime-zone/local-time info box, populated when Google renders one for the query. Always includes rawText plus structured sub-fields: time, date, dayOfWeek, timezoneAbbreviation (e.g. "EDT", or "GMT+9" for zones without a common abbreviation), and location. These are more reliably populated than the weather sub-fields since time-zone boxes use a simpler, consistent DOM shape in every capture tested.
currencyConverterBoxCurrency conversion box, populated for currency-conversion queries (e.g. "100 usd to eur"). Structured sub-fields: fromValue/fromUnit (input amount and currency code), toValue/toUnit (converted amount and currency code — toValue prefers Google's unrounded data-value over the rounded displayed text when both are present), exchangeRate, asOf (the rate's as-of timestamp text), and source. Any sub-field can independently be null if Google's markup for a given capture doesn't match the expected shape.
unitConverterBoxUnit conversion box, populated for unit-conversion queries (e.g. "5 km to miles"). Structured sub-fields: category (e.g. "Length"), fromValue/fromUnit, toValue/toUnit, and formula (the human-readable conversion formula text). Any sub-field can independently be null if Google's markup for a given capture doesn't match the expected shape.
jobsPackJob listing results (title, company, location, source) for job-search queries.
videoPack"Videos" carousel results (title, real YouTube URL, channel, duration, position) for video-intent queries (how-to/tutorial-style searches). Recovered from the same YouTube video cards (div.PmEWq / div[data-vid]) that organicResults already has to detect and exclude so they aren't miscounted as organic results — this field surfaces that already-located data instead of discarding it. channel and duration are independently null if a given card's markup doesn't expose them; title and url are always populated for every emitted card.
shoppingPackProduct results grid (title, price, original/"was" price, seller, rating, review count, position) for commercial queries (e.g. "buy running shoes"), gated on Google's own "More products"/"Shopping" [role="heading"] so it's never confused with unrelated carousels. No product URL is included: this grid is a JS-driven interactive "product viewer" widget with no real <a href> in the static HTML (navigation is injected by client-side JS on click), so a destination link genuinely isn't present to extract. originalPrice is only populated for cards actually on sale; rating/reviewCount are independently null if a card doesn't expose a rating.
stockBoxCurrently a known gap, always null. Real Google stock-price results use a different markup shape (a knowledge-panel/finance-card pattern) than the shared answer-box detector this actor uses for weather/time-zone. This is honestly disclosed as unimplemented against real markup rather than silently broken.
dictionaryBoxWord-definition box, populated for single-word or "define X"-style queries when Google renders one (e.g. "ubiquitous meaning"). Always includes rawText plus structured sub-fields: word (the plain headword, with Google's syllable-separator dots stripped), phonetic (e.g. "/yo͝oˈbikwədəs/"), partOfSpeech (e.g. "adjective"), and definitions — a list, one entry per sense in Google's own display order (more than one for multi-sense words), each with a definition string and an optional example usage sentence (present only when Google shows one for that sense). Any sub-field can independently be null/absent if Google's markup for a given capture doesn't match the expected shape — rawText is always the reliable fallback.

Not yet supported

The following Google SERP verticals are explicitly out of scope for this actor and are not extracted at all:

  • Images — Google serves base64 placeholder src values in the static HTML, not real image URLs. Extracting real images would need network interception or parsing a separate JS data callback — a fundamentally different approach from the rest of this actor.
  • Flights

Mobile support

Enabling Mobile Results swaps the user-agent and viewport to request Google's mobile SERP. People Also Ask, weather/time-zone answer boxes, and the local pack work the same way on mobile as desktop. Organic result extraction on mobile is experimental: Google's mobile page structure for organic results differs from desktop, so reliable extraction there isn't guaranteed. AI Overview parity on mobile is also unconfirmed. Treat mobile results as best-effort, especially for organicResults.

Reliability features

  • Parallel query processing — multi-query runs process several search terms at once instead of one after another, so a long keyword list finishes in a fraction of the time and one slow query no longer holds up the rest of the list.
  • site: results verified, not assumed — when you set the Site input, results are checked against that domain and any off-domain results Google slips in are dropped, so you never get contaminated data silently passed off as a domain-restricted search. This verification is specific to site: — the intitle:/intext:/inurl:/filetype: operators are sent to Google as-is and are not checked against the returned results, so Google can still include pages that don't actually match them.
  • Automatic reliability escalation — if the scraper detects it's being blocked, it automatically adapts and retries with a more robust approach mid-run instead of failing the query outright.
  • Broadened block/CAPTCHA detection — including reCAPTCHA challenge pages, not just generic error pages.
  • Per-query proxy rotation — each query gets a fresh proxy session.
  • Automatic proxy fallback, including Web Unblocker — if the default proxy setup gets blocked (for example on a hard CAPTCHA), the scraper automatically falls back through residential proxy and then Apify's Web Unblocker before giving up. This needs no configuration; leave Proxy Configuration unset and you get the whole chain.
  • Typed error rows per query — a query that errors out or that returns nothing because every fetch attempt was blocked produces a {"type": "error", "query": ..., "reason": ...} row in the dataset instead of aborting the whole run, so one bad query never takes down results for the rest.

Use cases

  • SEO monitoring — Track your website's ranking position for target keywords over time
  • Competitor analysis — See who ranks for the same keywords and what their listings look like
  • Market research — Discover what questions people ask about your industry (People Also Ask) and what Google's AI Overview says
  • Content ideas — Find related search queries to plan your content strategy
  • Local SEO — Pull local pack results for location-based queries
  • Lead generation — Find businesses and websites ranking for specific services

Supported countries

United States, United Kingdom, Canada, Australia, Germany, France, Spain, Italy, Brazil, Mexico, India, Japan, South Korea, Netherlands, Belgium, Austria, Switzerland, Sweden, Norway, Denmark, Finland, Poland, Czech Republic, Portugal, Ireland, New Zealand, South Africa, Singapore, Hong Kong, Taiwan, Philippines, Thailand, Indonesia, Malaysia, Vietnam, Argentina, Chile, Colombia, Peru, Turkey, Russia, Ukraine, Israel, UAE, Saudi Arabia, Egypt, Nigeria, Kenya

Export options

Results can be exported as JSON, CSV, Excel, HTML, XML, or RSS. Enable CSV Friendly Output to get one result per row — ideal for spreadsheet analysis.

CSV Friendly Output field reference

This is a lossy reshape, not just a re-formatting of the same data. With CSV Friendly Output enabled, each row is either an organic result or a paid result, and only the fields below are emitted:

typeFields on that row
organicquery, page, type, position, title, url, displayedUrl, snippet, siteLinks (site links flattened to "Title (url)", joined with |; omitted when the result has none)
paidquery, page, type, title, url

Everything else this actor extracts — aiOverview, knowledgePanel, peopleAlsoAsk, relatedQueries, localPack, topStories, sportsTable, and every Tier 2 opt-in box (weatherBox, timeZoneBox, currencyConverterBox, unitConverterBox, jobsPack, videoPack, shoppingPack, dictionaryBox, stockBox) — does not map onto a one-result-per-row shape and is not included anywhere in CSV mode's output. If you need any of that data, leave CSV Friendly Output off and use the default JSON/table export instead (see Output fields above for the full field list).

Integrations

Connect this scraper with your existing tools using the Apify API. Automate your workflow by integrating with Zapier, Make, Google Sheets, Slack, and more. Schedule runs to collect data on a regular basis.

FAQ

Is this an official Google product? No. This is an independent, third-party actor that scrapes publicly available Google Search result pages. It is not affiliated with, endorsed by, or sponsored by Google.

How fresh is the data? Every run fetches live results directly from Google at the moment the actor runs — there's no caching or stale data. Results reflect whatever Google is showing for that query right now.

How many results can I get? Up to 200 results per query across up to 10 pages of search results.

Does it work for non-English searches? Yes. Set the country and language to get localized results in any supported language. localPack's rating/reviewCount/address are parsed structurally from Google's markup rather than by matching English-phrased text, so parsing itself doesn't depend on the UI language — this works correctly against comma-decimal ratings (e.g. German "4,5"), Eastern/Extended Arabic-Indic digit review counts (e.g. Arabic "٣٬٦٧٩"), word-based review-count magnitude suffixes across several languages, and Hebrew's RTL-marked review-count/address rendering. As with any locale-parsing feature, an as-yet-unseen locale-specific number format could still surface a gap. resultsTotal uses the same locale-independent parsing approach (e.g. German "Ungefähr 104 Ergebnisse", Japanese "104 件の結果", Arabic "حوالي 15,900,000 نتيجة") and is reliably populated across locales — see the resultsTotal note in Output fields above.

Can I filter results by date? Yes, two ways: the Time Period option for relative ranges (past hour/day/week/month/year), or Date From / Date To for a custom absolute range. The two date fields must be set together.

Does it capture Google Ads / paid results? The scraper looks for paid ad slots on every request, but Google renders zero ads for a large share of queries in practice, so paidResults is often empty even for commercial search terms. Treat it as best-effort rather than a guaranteed field. All organic results, AI Overview, related queries, and People Also Ask data are unaffected by this.

I set a Site filter — why did I get fewer results than I asked for? Google only honors a site: restriction while it still has matching pages to show. Once it runs out, it quietly widens the search and starts returning results from completely unrelated domains, with nothing in the page marking them as such — this shows up most often from page 2 onwards. The scraper checks every result against the domain you asked for and drops the ones that don't belong (subdomains like en.wikipedia.org correctly count as matching wikipedia.org). So a smaller, genuinely domain-restricted result set is expected behavior, not a failure — it means Google ran out of pages on that site, and you're getting the accurate answer instead of padded, misleading data.

I set Word in Title / Word in Text / Word in URL / File Type — why do some results not actually match? Only the site: operator is verified against the returned results (see above). intitle:, intext:, inurl:, and filetype: are sent to Google as part of the query string as-is, but the scraper does not check the returned results against them the way it does for site:. Google itself doesn't strictly enforce these operators either — it can and does slip in results that don't actually satisfy them, and those pass through unfiltered. Treat these as best-effort relevance hints rather than a hard guarantee, and spot-check results if exact compliance matters for your use case.

Why is aiOverview empty for my query? By default, this is a known limitation rather than a reflection of whether Google showed an AI Overview for your query. Google injects the AI Overview content into the page via client-side JavaScript after load, so it's only present when the scraper renders the page in JavaScript mode — never in the actor's default, faster fetch mode. The actor only switches to the slower mode after repeated blocks, so on a normal, unblocked run aiOverview will be null even for queries where Google genuinely rendered one.

How do I try to get aiOverview populated? Enable the Force AI Overview capture (forceAiOverviewCapture) input option. When set, every fetch for that run renders the page in JavaScript mode and the actor explicitly waits for Google's AI Overview container to finish loading before reading the page, instead of only doing so as an emergency fallback after repeated blocks. This makes each request slower and more resource-intensive, and it's best-effort, not a guarantee — Google may genuinely not show an AI Overview for a given query.

Tip: for the most reliable AI Overview capture, set Proxy Configuration to the Web Unblocker group alongside this flag — other proxy options are less likely to succeed for this specific feature. It's still best-effort: Google may genuinely not render an AI Overview for a given query. With the flag on, if that fetch fails, the actor automatically falls back to the default fetch mode so organic/other results are still returned for that query — just with aiOverview: null.

Does the url field in organic results point straight to the destination site? No — it's Google's own redirect link (https://www.google.com/goto?url=<opaque token>), and the token isn't a decodable form of the destination URL. Use displayedUrl for the actual site domain (see the caveat about social-media citation cards in the Output fields table above).