Google AI Overview Scraper
Pricing
from $3.90 / 1,000 results
Google AI Overview Scraper
A robust, high-performance utility designed for developer automation, data integration, and AI training. Features built-in captcha bypass, headful/headless browser execution, and proxy support to scrape Google data seamlessly, reliably, and at scale.
Pricing
from $3.90 / 1,000 results
Rating
0.0
(0)
Developer
Coding Frontned
Maintained by CommunityActor stats
0
Bookmarked
155
Total users
5
Monthly active users
10 days ago
Last modified
Categories
Share
What this Actor does
This Actor collects AI-generated answer text and visible citations from Google Search. It emits a dataset row only when Google serves a recognizable AI Overview with substantial answer text. A normal search page without an overview is reported in the OUTPUT run summary and does not create a placeholder row.
AI Overview availability depends on the query, country, language, account/session context, Google experiments, and current serving decisions. The Actor reports those boundaries instead of treating an unavailable overview as scraped content.
Features
- Processes up to 20 normalized, deduplicated queries.
- Supports
gl,hl, bounded retries, request timeouts, respectful request delay, and standard Apify proxy configuration. - Captures visible answer text, related questions, citation titles, displayed URLs, favicons, and resolved external citation URLs when Google exposes redirect wrappers.
- Preserves
googleRedirectUrlwhen a citation destination cannot be resolved. - Keeps operational diagnostics in the
OUTPUTkey-value record rather than adding them to normal dataset rows. - Detects Google challenge or malformed-page boundaries and fails the run atomically instead of publishing misleading rows.
Input
| Field | Type | Default | Description |
|---|---|---|---|
queries | string[] | required | One to 20 non-empty queries, each at most 300 characters. |
maxItems | integer | 10 | Maximum number of input queries to process. |
gl | string | us | Two-letter Google country code. |
hl | string | en | Language code, optionally followed by a country code. |
maxRequestRetries | integer | 1 | Zero to two retries per query. |
requestTimeoutSecs | integer | 60 | Per-request timeout from 15 to 90 seconds. |
requestDelayMs | integer | 0 | Optional delay from 0 to 5,000 milliseconds before each query. |
proxyConfiguration | object | Google SERP proxy | Standard Apify proxy settings. GOOGLE_SERP is recommended for Google result pages. |
Default request
{"queries": ["What is quantum computing?"],"maxItems": 1}
Locale and developer controls
{"queries": ["benefits of green tea", "how does photosynthesis work"],"maxItems": 2,"gl": "us","hl": "en","maxRequestRetries": 1,"requestTimeoutSecs": 60,"requestDelayMs": 500,"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["GOOGLE_SERP"]}}
For a local diagnostic request, proxyConfiguration: { "useApifyProxy": false } uses the local network. A custom proxyUrls list must be paired with useApifyProxy: false; Apify Proxy and custom proxy URLs cannot be combined.
Output
The dataset contains only genuine overview records:
{"query": "What is quantum computing?","aiOverviewText": "Quantum computing uses quantum mechanical phenomena to process information...","sources": [{"title": "Introduction to quantum computing","url": "https://example.org/quantum","googleRedirectUrl": "https://www.google.com/goto?url=https%3A%2F%2Fexample.org%2Fquantum"}],"relatedQuestions": ["How does quantum computing work?"],"searchUrl": "https://www.google.com/search?q=What+is+quantum+computing%3F&gl=us&hl=en","scrapedAt": "2026-01-01T00:00:00.000Z"}
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
The OUTPUT key-value record contains success, item counts, per-query statuses, failures, and elapsed time. validEmpty: true means every query completed without transport or page-boundary failure but Google served no overview for the tested queries.
Limitations and responsible use
Google may return different result markup or no AI Overview for the same query over time. CAPTCHA, unusual-traffic pages, empty responses, and unrecognized markup are reported as failures. The Actor does not spoof browser fingerprints, bypass CAPTCHAs, or store raw HTML. It uses bounded public requests and standard proxy settings.
Local development
npm installnpm testnpm run lintnpm run validatenpx --no-install apify validate-schemanpx --no-install apify run --purge --input-file ./qa-inputs/green-tea.json