Google AI Overview Scraper avatar

Google AI Overview Scraper

Pricing

from $3.90 / 1,000 results

Go to Apify Store
Google AI Overview Scraper

Google AI Overview Scraper

A robust, high-performance utility designed for developer automation, data integration, and AI training. Features built-in captcha bypass, headful/headless browser execution, and proxy support to scrape Google data seamlessly, reliably, and at scale.

Pricing

from $3.90 / 1,000 results

Rating

0.0

(0)

Developer

Coding Frontned

Coding Frontned

Maintained by Community

Actor stats

0

Bookmarked

155

Total users

5

Monthly active users

10 days ago

Last modified

Share

What this Actor does

This Actor collects AI-generated answer text and visible citations from Google Search. It emits a dataset row only when Google serves a recognizable AI Overview with substantial answer text. A normal search page without an overview is reported in the OUTPUT run summary and does not create a placeholder row.

AI Overview availability depends on the query, country, language, account/session context, Google experiments, and current serving decisions. The Actor reports those boundaries instead of treating an unavailable overview as scraped content.

Features

  • Processes up to 20 normalized, deduplicated queries.
  • Supports gl, hl, bounded retries, request timeouts, respectful request delay, and standard Apify proxy configuration.
  • Captures visible answer text, related questions, citation titles, displayed URLs, favicons, and resolved external citation URLs when Google exposes redirect wrappers.
  • Preserves googleRedirectUrl when a citation destination cannot be resolved.
  • Keeps operational diagnostics in the OUTPUT key-value record rather than adding them to normal dataset rows.
  • Detects Google challenge or malformed-page boundaries and fails the run atomically instead of publishing misleading rows.

Input

FieldTypeDefaultDescription
queriesstring[]requiredOne to 20 non-empty queries, each at most 300 characters.
maxItemsinteger10Maximum number of input queries to process.
glstringusTwo-letter Google country code.
hlstringenLanguage code, optionally followed by a country code.
maxRequestRetriesinteger1Zero to two retries per query.
requestTimeoutSecsinteger60Per-request timeout from 15 to 90 seconds.
requestDelayMsinteger0Optional delay from 0 to 5,000 milliseconds before each query.
proxyConfigurationobjectGoogle SERP proxyStandard Apify proxy settings. GOOGLE_SERP is recommended for Google result pages.

Default request

{
"queries": ["What is quantum computing?"],
"maxItems": 1
}

Locale and developer controls

{
"queries": ["benefits of green tea", "how does photosynthesis work"],
"maxItems": 2,
"gl": "us",
"hl": "en",
"maxRequestRetries": 1,
"requestTimeoutSecs": 60,
"requestDelayMs": 500,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["GOOGLE_SERP"]
}
}

For a local diagnostic request, proxyConfiguration: { "useApifyProxy": false } uses the local network. A custom proxyUrls list must be paired with useApifyProxy: false; Apify Proxy and custom proxy URLs cannot be combined.

Output

The dataset contains only genuine overview records:

{
"query": "What is quantum computing?",
"aiOverviewText": "Quantum computing uses quantum mechanical phenomena to process information...",
"sources": [
{
"title": "Introduction to quantum computing",
"url": "https://example.org/quantum",
"googleRedirectUrl": "https://www.google.com/goto?url=https%3A%2F%2Fexample.org%2Fquantum"
}
],
"relatedQuestions": ["How does quantum computing work?"],
"searchUrl": "https://www.google.com/search?q=What+is+quantum+computing%3F&gl=us&hl=en",
"scrapedAt": "2026-01-01T00:00:00.000Z"
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

The OUTPUT key-value record contains success, item counts, per-query statuses, failures, and elapsed time. validEmpty: true means every query completed without transport or page-boundary failure but Google served no overview for the tested queries.

Limitations and responsible use

Google may return different result markup or no AI Overview for the same query over time. CAPTCHA, unusual-traffic pages, empty responses, and unrecognized markup are reported as failures. The Actor does not spoof browser fingerprints, bypass CAPTCHAs, or store raw HTML. It uses bounded public requests and standard proxy settings.

Local development

npm install
npm test
npm run lint
npm run validate
npx --no-install apify validate-schema
npx --no-install apify run --purge --input-file ./qa-inputs/green-tea.json